Skip to content
Made with Jev

What Jev decides in a mobile agent — and what stays on vision

How CoreAutomata's phone agent splits work between Jev, TypeSafe's System One model, and our vision model — and the one rule we had to fix along the way.

CoreAutomata Team · · 6 min read

Cover: What Jev decides in a mobile agent — and what stays on vision
On this page
  1. Why a phone agent needs two kinds of model
  2. What Jev is, and what it isn't
  3. The decisions we hand to Jev
  4. What stays on vision
  5. The rule we had to fix: delegate means ask, not tap
  6. Failing open
  7. What we'll publish next
  8. Try it on your own phone

A phone agent makes two very different kinds of decision. Some need to look: where exactly is the button on a screen full of pixels, what does this chart say, is this a game board or a login form. Others only need to choose: given the controls this screen says it has, which one moves the task forward?

CoreAutomata’s agent now hands the second kind to Jev, TypeSafe’s System One model, and keeps our vision model for the first. This post walks through exactly where that line sits in the product today, what happens when Jev is unsure, and the rule we had to fix before we were happy calling it “made with Jev”.

Why a phone agent needs two kinds of model

Every step of a CoreAutomata run starts with the same two things: a screenshot of the live phone and its accessibility tree — the list of controls the app exposes, with their labels, roles, and whether they can be tapped, typed into, or scrolled.

When the tree is rich, the next step is often a multiple-choice question. The screen says it has Orders, Returns, Settings and a search box; the task is “export last month’s orders”. A model that can read that list and pick Orders, with an honest sense of how sure it is, doesn’t need to look at a single pixel.

When the tree is thin — a game, a canvas, a custom-drawn chart, an app that labels nothing — there is nothing to choose from, and the only way forward is to look. That is what vision models are for.

What Jev is, and what it isn’t

Jev doesn’t write text. You give it some state and a set of named options, and it returns one of those options with a probability for each. It can also answer yes/no questions with a calibrated score. That shape is what makes it a good fit for the choose half of a phone agent: the answer is always one of the things we offered, never an invented control.

It also means Jev can’t do several things an agent needs. It doesn’t read screenshots. It doesn’t plan a ten-step task. It doesn’t count totals, compare dates, or write a reply to a message. Those stay with our vision model.

The decisions we hand to Jev

These are the places CoreAutomata calls Jev today.

  • Which control to tap next. For goal-driven tasks and guided test authoring, we build a menu from the labelled, enabled controls on the live screen — “Tap Orders”, “Type ‘alice’ into Username”, “Swipe up on Results” — plus one extra option: delegate. Jev picks one. We only act on the pick when Jev is confident and clearly ahead of its runner-up.
  • Whether the task is already done. Jev’s menu has no “finish” option, so the same request also asks whether the goal already looks satisfied. Jev acts only on a clear “not yet”. Anything else goes to our larger model, which decides whether to stop.
  • Which element a test step means. When a saved test step says “tap Checkout” and our deterministic matcher can’t find one clear winner, Jev gets the candidates first. A confident pick still goes through the same conflict checks as before; anything else falls through to the vision grounder.
  • Whether two screens are the same. During auto discovery, Jev tells us when a “new” screen is really one we’ve already mapped with different list content, so the crawl doesn’t waste steps on repeats.
  • Whether a check in a run report passed. For plain-language criteria in a journey report, Jev answers “does the run clearly show this was met?” Criteria that involve counts, totals, amounts or dates skip Jev and go straight to the larger model.

What stays on vision

  • Screens with little readable text. If the accessibility tree is sparse, we don’t ask Jev at all. Games, canvases and custom-drawn UI go straight to vision.
  • Telling twins apart. If Jev picks “Add” and the screen has two controls labelled “Add”, we take a screenshot and ask the vision grounder where that control is — grounded on Jev’s pick, not on the overall goal. Vision can move the tap to the right “Add”. It never swaps Jev’s pick for an unrelated pixel tap, which keeps saved test steps readable (tap 'Add', not tap (412, 1180)).
  • Anything Jev delegates. When Jev chooses delegate, isn’t confident enough, or the task might already be done, our vision model runs the step exactly as it did before Jev existed.
  • Planning and writing. Breaking a task into sub-goals, drafting a reply, summarising what a screen says — all vision-model work.

The rule we had to fix: delegate means ask, not tap

Our first version ran Jev and the vision grounder side by side on every step and merged the answers. It looked efficient on paper. In practice it had a flaw: when Jev said delegate — its way of saying “I can’t answer this from the labels” — the merge step still had a vision coordinate in hand, so it tapped wherever vision pointed for the overall goal.

That quietly turned Jev’s most useful signal into a blind tap. Worse, because the larger model never ran, the agent could never choose to type, go back, or finish — so guided authoring could keep tapping after the task was already done.

We changed the rule. A delegate now always means the larger model decides the step. Vision only runs alongside Jev when Jev’s pick is ambiguous, and then only to confirm where that pick is. Confident, unique picks skip the screenshot round-trip entirely — which is where the savings were supposed to come from in the first place.

Failing open

Jev sits in front of paths that already worked, so every call fails open. No key, a timeout, an error, or an unusable answer all return “no decision”, and the step runs on our vision model as before. Each call has one hard time budget with no silent retries stacked on top, and the connection is reused rather than rebuilt per call. Jev can make a run faster and cheaper; it is never allowed to make one fail.

What we’ll publish next

We haven’t published speed or cost numbers yet, on purpose. The next post will take a few real tasks on a paired Android phone and report, per run, how many decisions Jev made versus our vision model, the time each took, and what the run cost — measured from our own traces, with the date and the task named.

Try it on your own phone

You don’t need a Jev key or a model of your own: CoreAutomata hosts all of it. Pair an Android phone or an iPhone by QR, describe a task in plain English, and watch which steps Jev picks and which go to vision — every step is recorded with its screenshot.

Tags: jev · typesafe · system-one · mobile-agents · architecture