Back to AI Research

AI Research

PhoneCLI: From App Interfaces to Callable Commands... | AI Research

Key Takeaways

  • What the paper is about Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model...
  • Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action.
  • It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation is static, ordered, and endlessly repeated.
  • We present PhoneCLI, which compiles an app's GUI navigation into callable commands, without any app-internal API, runtime instrumentation, or model training.
  • On AndroidLab, PhoneCLI improves the task success rate while reducing steps and token consumption, and it transfers to AndroidWorld's official M3A agent with consistent efficiency gains.
Paper AbstractExpand

Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation is static, ordered, and endlessly repeated. We present PhoneCLI, which compiles an app's GUI navigation into callable commands, without any app-internal API, runtime instrumentation, or model training. Offline, PhoneCLI explores a target app from the outside and distills its screens, interactive elements, and navigation edges into a semantically annotated map; each screen yields one deterministic command: a replay sequence that reaches it. Online, the agent selects a command, verifies it before execution, and then executes it deterministically in sub-second time at zero VLM cost; open-ended interaction and every failure of the compiled path fall back to the embedded VLM interpreter, exactly the pure VLM agent, so compilation can only help. On AndroidLab, PhoneCLI improves the task success rate while reducing steps and token consumption, and it transfers to AndroidWorld's official M3A agent with consistent efficiency gains. What PhoneCLI compiles is the app's navigation rather than one run, so it serves new tasks, not only repeated ones.

What the paper is about

Mobile GUI agents operate through a perception--action loop: at each step they screenshot the device, invoke a vision--language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation---and everyday navigation is static, ordered, and endlessly repeated. We present PhoneCLI, which compiles an app's GUI navigation into callable commands, without any app-internal API, runtime instrumentation, or model training. Offline, PhoneCLI explores a target app from the outside and distills its screens, interactive elements, and navigation edges into a semantically annotated map; each screen yields one deterministic command: a replay sequence that reaches it. Online, the agent selects a command, verifies it before execution, and then executes it deterministically in sub-second time at zero VLM cost; open-ended interaction and every failure of the compiled path fall back to the embedded VLM interpreter, exactly the pure VLM agent, so compilation can only help. On AndroidLab, PhoneCLI improves the task success rate while reducing steps and token consumption, and it transfers to AndroidWorld's official M3A agent with consistent efficiency gains. What PhoneCLI compiles is the app's navigation rather than one run, so it serves new tasks, not only repeated ones. The same ai evaluation question is explored in MTVA-Bench, which adds a research perspective.

What it covers

PhoneCLI: From App Interfaces to Callable Commands for Mobile Agents Yangqin Jiang Lingrui Xu Chao HuangThe University of Hong Kong † † thanks: Chao Huang is the Corresponding Author. Email: {mrjiangyq99, lingruixu.db, chaohuang75}@gmail.com ˜ Github Repo: https://github.com/HKUDS/OpenPhone Abstract Mobile GUI agents operate through a perception–action loop: at each step they screenshot the device, invoke a vision–language model (VLM), and emit an action. It is slow, costly, and brittle, yet most of what it does is navigation—and everyday navigation is static, ordered, and endlessly repeated. We present PhoneCLI, which compiles an app’s GUI navigation into callable commands, without any app-internal API, runtime instrumentation, or model training. Offline, PhoneCLI explores a target app from the outside and distills its screens, interactive elements, and navigation edges into a semantically annotated map; each screen yields one deterministic command: a replay sequence that reaches it. Online, the agent selects a command, verifies it before execution, and then executes it deterministically in sub-second time at zero VLM cost; open-ended interaction and every failure of the compiled path fall back to the embedded VLM interpreter, exactly the pure VLM agent, so compilation can only help. On AndroidLab, PhoneCLI improves the task success rate while reducing steps and token consumption, and it transfers to AndroidWorld’s official M3A agent with consistent efficiency gains. What PhoneCLI compiles is the app’s navigation rather than one run, so it serves new tasks, not only repeated ones. 1 Introduction GUI agents for mobile devices largely operate through a perception–action loop: at each step, the agent captures the screen, invokes a vision–language model (VLM) to reason over the visual information, and emits an action at either the index level or in coordinate form. The agent’s competence is acquired through training on GUI corpora and/or via fine-tuning and reinforcement learning [ 1 , 2 , 3 , 4 ] . This paradigm is confronted with three compounding challenges. (I) Closedness. Desktop computer-using agents can operate via rich programmatic surfaces— e.g. , CLI tools, APIs, or even executable code [ 5 , 6 , 7 ] —because the desktop ecosystem provides open interfaces. By contrast, mobile apps offer no comparable privilege. They function as black boxes with no client-side, third-party programmatic interface, and even modern mobile frameworks typically only allow composition with system-level commands [ 8 ] , leaving the internal state and logic of each app inaccessible. Consequently, the agent is restricted to a screenshot-based loop, which is slow (on the order of seconds per step), costly (requiring one API call per step), and brittle, often leading to hallucinated coordinates and navigation failures. (II) Repetitiveness. Nevertheless, the navigation induced by the loop targets interfaces whose structure is largely known a priori. Everyday mobile app usage is highly repetitive and follows consistent, ordered patterns [ 9 , 10 ] : for instance, opening a settings page, selecting a tab, and locating a specific item. Such behavior typically traces a small set of static paths repeatedly. Unlike free-form interaction, the navigation structure can therefore be enumerated offline. This motivates the core idea of this work: pre-build the repeated navigation once and replay it deterministically thereafter. (III) Generalization. The alternative route—training—stores capability in the model weights. However, each new app typically requires additional data collection and/or re-training, which exposes a major bottleneck studied in recent work on RL generalization [ 11 ] and on adaptation to previously unseen apps [ 12 ] . Re-training is both costly and inflexible: a model that has never encountered an app has no direct shortcut for navigating it. Given the scale of modern app stores, with millions of apps, per-app training is infeasible in principle. Existing remedies either improve the loop from within—by pairing stronger VLMs with grounding and memory scaffolding [ 13 , 14 ] —or incur the training cost to sharpen the agent [ 2 ] . However, neither approach eliminates the repeated navigation bottleneck itself. We propose PhoneCLI , which takes a third route: compile an app’s GUI navigation into callable commands . Offline , PhoneCLI explores each target app once from the outside, without any app-internal API and without instrumentation. It distills the app’s screens, UI elements, and navigation edges into a structured map, and compiles each screen into a deterministic command— i.e. , a replayable action sequence that reliably reaches that screen. This offline procedure does not involve training or model updates, so integrating a new app takes only minutes. Online , the agent maps each task to a command, verifies the selection, and replays it deterministically in sub-second time with zero VLM calls. Only open-ended interactions not covered by the command table are handled by the VLM, which serves as a graceful fallback. Recent work shares parts of this picture. PreAct [ 15 ] compiles agent trajectories online to replay the same task. UI-KOBE [ 16 ] enables runtime following of an explored UI graph. Trajectory-mining approaches recycle explored paths back into the VLM as textual hints [ 17 , 18 ] , meaning that reuse still incurs perception costs. PhoneCLI differs along the decisive axis: it compiles offline, at the granularity of navigation commands, and then executes them deterministically. As a result, the compiled knowledge supports unseen tasks rather than only repeated ones, and its reuse requires no perception at runtime. In short, PreAct compiles what an agent has done, whereas PhoneCLI compiles what the app is. We discuss our contributions are threefold:

• GUI-to-CLI compilation. We present a compilation pipeline that transforms a black-box app—lacking any app-internal API, instrumentation, and training—into a structured app map and a catalog of deterministic callable commands, one per screen. This pipeline provides the agent with a programmatic surface that the app itself does not expose, enabling new apps to be integrated within minutes without any model update.

• Reliable command invocation. We design an invocation mechanism based on two-phase routing that selects and verifies a command prior to execution. Deterministic replay places the agent in the target screen in sub-second time at zero VLM cost, while any failure falls back gracefully to the embedded VLM interpreter. By removing the most frequently repeated aspect of mobile app usage—navigation—from the VLM loop, PhoneCLI preserves the performance of its backbone without degradation.

• Dual-benchmark evidence. On AndroidLab [ 19 ] , PhoneCLI achieves state-of-the-art performance with zero training, while reducing the step and token cost of successful tasks. On AndroidWorld [ 20 ] , it transfers to the official agent with measurable efficiency gains. A controlled three-way ablation further attributes the improvements to the map information and deterministic replay, independently. 2 Methodology Figure 1: Overall framework of the proposed PhoneCLI. 2.1 Design Overview Motivation and design rationale. As detailed in Section 1 , our design rests on three observations: (I) closedness , mobile apps expose no programmatic surface, unlike the CLI tools and APIs available on desktop, which confines mobile agents to a slow, costly, and brittle screenshot-perception loop; (II) repetitiveness , everyday app usage is highly repetitive and ordered, and the navigation skeleton that connects screens changes far more slowly than app content, so a path distilled once is replayed many times; and (III) generalization , training-based agents are prone to distribution shift on unseen apps and costly to re-train, whereas on-demand exploration integrates a new app in minutes. This motivates our central design principle: build once what is static, and handle on the fly what is not . Navigation is static, frequent, and pre-enumerable, so we distill it offline into a table of callable commands and reuse it across tasks; open-ended interaction—form filling, dynamic content, unexpected dialogs—is variable and cannot be exhaustively enumerated, so it is handled by the VLM at runtime. Such content—feeds, recommendations, merchant lists—changes on every load; the navigation skeleton does not, and only the skeleton is compiled. We refer to the offline distillation as compilation and to the runtime handling as interpretation —a division of labor analogous to mixed-mode execution in programming languages, where frequently used code is compiled ahead of time and the remainder is interpreted on demand. Accordingly, PhoneCLI comprises three stages:

• Offline compilation (Stage 1, Sec. 2.2 ). An automated exploration traverses the app from a clean state and compiles its navigation skeleton into a semantically annotated map plus a catalog of commands, turning an interface that can be operated only manually into one that can be invoked programmatically.

• Online invocation (Stage 2, Sec. 2.3 ). A router maps a task to a compiled command, the match is confirmed, and the command is replayed deterministically—the analog of invoking a compiled function, with a guard on either side of the call.

• Runtime interpretation (Stage 3, Sec. 2.4 ). Everything outside the command table is handled by the VLM, which receives every failure of the compiled path as a graceful degradation. Formalization. Formally, an app map is a structure 𝒢 = ( S , E , λ , M ) \mathcal{G}=(S,E,\lambda,M) , where S S is the set of discovered screens; E E is the set of interactive elements, each annotated with normalized coordinates; λ : E ⇀ S × Alias ∗ \lambda:E\rightharpoonup S\times\mathrm{Alias}^{*} is a partial function that records navigation by mapping an element to the screen reached upon its activation together with its semantic aliases, so that a navigation edge is a triple ( s , e , s ′ ) (s,e,s^{\prime}) with e e an element of s s ; and M M is a set of commands , one per screen. Each command c = [ a 1 , … , a k ] c=[a_{1},\dots,a_{k}] is a deterministic action sequence that navigates to its target screen. Stage 2 consumes a catalog 𝒞 \mathcal{C} derived from 𝒢 \mathcal{G} : each entry pairs a stable identifier with the natural language description and target of a command, and carries neither coordinates nor action steps, so that routing operates on symbols rather than on plans. A router r ⁡ ( T , 𝒞 ) ↦ ( d , 𝑖𝑑 , 𝑎𝑛𝑠 ) r(T,\mathcal{C})\mapsto(d,\mathit{id},\mathit{ans}) maps a task description T T to a decision d d , the identifier of a selected command when one applies, and an answer string 𝑎𝑛𝑠 \mathit{ans} when d = FINISH d=\texttt{FINISH} ; the identifier resolves to a command c ∈ M c\in M . The decision takes one of four values: d = OP d=\texttt{OP} , the command alone completes the task and is replayed with the answerability check; d = MACRO_VLM d=\texttt{MACRO_VLM} , the compiled macro performs the required navigation and then transfers control to the VLM, and we call a screen’s replay sequence its macro; d = NEED_VLM d=\texttt{NEED_VLM} , no applicable command exists and the VLM explores from the current state; and d = FINISH d=\texttt{FINISH} , the task can be answered without device interaction. Design properties. (I) Compile once, reuse many. The offline cost is amortized across all subsequent tasks and runs, and navigation no longer consumes the VLM budget. (II) Guaranteed floor. Every failure path—no matching command, a rejected verification, or a landing mismatch—resolves to the pure VLM agent, so the success rate is never below that of the VLM-only baseline; what the compiled path adds is a bounded number of routing calls, not a new failure mode. (III) Model-agnostic knowledge acquisition. The command table is constructed without training or model updates; a new app is integrated in minutes by re-running the offline stage. 2.2 Offline Compilation: Building App Maps The first stage constructs the app map 𝒢 = ( S , E , λ , M ) \mathcal{G}=(S,E,\lambda,M) from a black-box app in two steps: structural exploration recovers the app’s navigation skeleton, and semantic enrichment renders that skeleton interpretable to a language model. Our exploration inherits systematic UI traversal from Android testing crawlers [ 21 , 22 ] , but targets reusable navigation rather than test coverage. Unlike trajectory-based methods that inject mined knowledge back into the VLM as text [ 17 , 18 ] , the explored structure is compiled into deterministic, executable commands. Structural exploration. Starting from a clean launch state, the crawler traverses the app in breadth-first order. At each screen it reads the UI hierarchy, records the interactive elements it finds, activates them one by one, and observes the screens they reach, thereby assembling the screens, elements, and navigation edges of 𝒢 \mathcal{G} . Scrollable screens are explored over multiple scroll pages so that content below the fold is not omitted. The walk is bounded by depth and screen budgets chosen to cover the navigation core of typical apps. Screen deduplication is computed from raw elements before semantic annotation, so that persistent bottom navigation bars shared across tabs do not dominate the signature and collapse distinct tabs into a single screen. Semantic enrichment. The raw graph captures syntax—text and coordinates—but not semantics: what an element is, or what a screen is for. Because both routing (Stage 2) and verification depend on semantic understanding, a language model annotates the graph. Each element is classified as STABLE (present across app versions) or DYNAMIC ; assigned aliases with a semantic type ( button , input , label , etc.); and each screen is given a natural language description. Compilation to commands. From the enriched map, the compiler derives one command per screen: the action sequence that reaches it from a clean launch—typically force_stop , launch , followed by a sequence of taps and swipes with fixed coordinates and inter-step waits. A command is executable rather than advisory: replaying it reproduces the recorded trajectory exactly, with no model in the loop. Because the map is obtained purely by exploration and LLM annotation, it can also be rebuilt whenever an app changes. An excerpt of the resulting app map is shown in Figure 6 (Appendix A.2 ). 2.3 Online Invocation: Routing and Replaying Commands Stage 2 is the online counterpart of the compiled command table: given a task, it decides whether a compiled command applies and, if so, executes it reliably. It is deliberately conservative—whenever the mapping is uncertain, the command is declined and the task proceeds to Stage 3. Command routing. Routing proceeds in two phases, each implemented by a lightweight LLM call. Phase I (selection) presents the task together with a compact command catalog and asks the router to choose among the four decisions defined in Section 2.1 . The catalog is formatted as a prefix tree-commands that share a navigation path are merged at their common prefix and expanded by indentation—so that it remains compact, and it is deliberately symbolic: it carries element paths, semantic tags, and stable operation identifiers, but neither coordinates nor action steps, which keeps it short and prevents the router from inventing device-level actions. A catalog excerpt, together with the router’s output format and an example decision, is shown in Figure 7 (Appendix A.3 ). Phase II (verification) inspects the selected command before any execution: the router judges whether the command’s target screen, described in natural language, is semantically consistent with the task. If it is not, the command is rejected and the task falls back to the VLM, since replaying a plausible but incorrect command would land the VLM on an irrelevant screen and waste the rounds spent recovering. Deterministic replay and completion. A selected identifier resolves to the command compiled from the map—the model chooses what to do, while the program retains how —so replay requires no perception, incurs no model call, and cannot hallucinate coordinates. Execution is therefore deterministic and completes in sub-second time, in contrast to screen-by-screen navigation, where every step costs seconds and one API call. This is where compilation pays off: navigation, the most repeated part of app use, is removed from the VLM loop entirely. A replayed command is then completed in one of two ways. If the command alone solves the task ( OP ), the agent performs an answerability check on the landed screen and finishes with that answer if it is available; this check differs in kind from the Phase II verification, which asks whether the command’s destination fits the task, whereas this one asks whether the reached screen actually yields the answer. If the answer is not available, the command is demoted to the MACRO_VLM path, so that a command which navigated correctly but cannot answer on its own still saves the VLM the navigation. To avoid a cold start on an unfamiliar screen, the handoff injects a landing hint: the target screen’s natural language description and the list of its interactive elements, together with an explicit note that the command has only navigated to this screen and that the interaction remains to be performed. The hint is cheap—it is drawn from the app map—and turns the apparent teleportation into a warm start. Landing check. Replay is not trusted blindly. After execution, the current UI hierarchy is matched against the recorded screens; if the agent has demonstrably landed on a screen unrelated to the command’s target, the app is restarted to a clean state before the VLM takes over, so that a mismatch never silently misleads it. The command is thus verified before replay and checked afterward, which makes the invocation path self-checking and routes every failure to Stage 3. The full invocation path is summarized in Algorithm 1 . Algorithm 1 PhoneCLI: executing one task Input: Task T T , app map 𝒢 = ( S , E , λ , M ) \mathcal{G}=(S,E,\lambda,M) , VLM Output: Task outcome (graded success or failure) 1 𝒞 ← \mathcal{C}\leftarrow catalog of 𝒢 \mathcal{G} ; // Stage 1: symbolic entries, built once per app 2 ( d , 𝑖𝑑 , 𝑎𝑛𝑠 ) ← r ⁡ ( T , 𝒞 ) (d,\mathit{id},\mathit{ans})\leftarrow r(T,\mathcal{C}) ; // Stage 2, Phase I: selection 3 switch d d do 4 case FINISH do 5 return 𝑎𝑛𝑠 \mathit{ans} // answerable without device interaction 6 end case 7 case NEED_VLM do 8 return Interpret (T) // no command applies; continue from the current state 9 end case 10 case OP / MACRO_VLM do 11 c ← Resolve ​ ( 𝑖𝑑 ) c\leftarrow\textsc{Resolve}(\mathit{id}) ; // identifier → \to recorded action sequence 12 if ¬ \neg Verify ( c , T c,T ) then 13 return Interpret (T) // Phase II rejects; current state retained 14 Restart (app); Replay ( c c ); // clean state; deterministic, zero VLM calls 15 if ¬ \neg LandingCheck ( c c ) then 16 Restart (app); return Interpret (T) // clean-state fallback 17 if d = OP d=\texttt{OP} ∧ \land answerable on the landed screen then 18 return the answer as 𝑎𝑛𝑠 \mathit{ans} // command alone completes the task 19 h ← h\leftarrow Hint ( c c ); return Interpret (T, h) // OP demoted, or MACRO_VLM; Stage 3 20 end case 21 end switch 2.4 Runtime Interpretation: VLM Fallback and Graceful Degradation Why interpretation is needed. Compilation is inherently partial: the map covers only what exploration could observe, and open-ended interaction cannot be enumerated in advance. Stage 3 therefore interprets what was not compiled. The VLM perceives the screen and acts step by step, following the standard perception—action loop of GUI agents [ 13 , 14 ] ; we write Interpret ​ ( T , h ) \textsc{Interpret}(T,h) for this loop, where the optional h h is the landing hint injected after a replayed navigation. Memory and recovery. The interpreter maintains a structured state across rounds. Each VLM response ends with a STATE_ASSESSMENT field, and past assessments are injected into the next prompt in chronological order, giving the agent an explicit memory of what it has tried and where it is stuck. On top of this, the loop detects unproductive behavior—repeated scrolling or waiting without screen progress—and injects targeted hints ( e.g. , try the search entry; close the ad overlay). When an action fails at execution time, the error is recorded and re-presented in the next round with a corrective hint. These mechanisms turn the interpreter from a stateless screen actor into a small but self-correcting agent. Graceful degradation. The degradation rule is uniform: whenever the compiled path is unavailable, the system runs the pure VLM agent, continuing from the current state, or from a restarted clean state when the landing check has failed. Stage 3 is thus a co-routine rather than a last resort—it is entered by design on every failure of the compiled path, and the interpreter it supplies is exactly the agent that would otherwise have run alone. 3 Evaluation 3.1 Experimental Setup Benchmarks. We evaluate on two Android agent benchmarks: AndroidLab [ 19 ] (138 tasks, 9 apps) and AndroidWorld [ 20 ] (116 tasks, 20 apps, open-world with cross-app workflows), where we compare against its official M3A agent. Backbone Models. We use three backbone models — Qwen3.7-Plus [ 23 ] , GLM-4.6V [ 24 ] , and Kimi-K3 [ 25 ] . AndroidWorld experiments use Qwen3.7-Plus for both the agent and the baseline, keeping the comparison model-controlled. Baseline methods. Comparisons use two primary model groups: (1) General-purpose vision-capable LLMs : closed-source models (Qwen3.7-Plus [ 23 ] , Gemini-2.5-Pro [ 26 ] , GPT-4o [ 27 ] , Claude-Sonnet 4) and open-weight models (Kimi-K3 [ 25 ] , GLM-4.6V [ 24 ] ) (2) GUI-specialized/fine-tuned models : AutoGLM-Phone, AutoGLM-Mobile [ 2 ] , MobileUse [ 3 ] , UI-Genie-Agent [ 28 ] , UI-Tars-1.5 [ 29 ] , V-Droid [ 30 ] , and AutoGLM-2024-10 [ 1 ] . More details of the experimental setup are reported in Appendix A.1 . 3.2 Overall Performance PhoneCLI is an enhancement layer attached to any GUI agent rather than a standalone backbone, so the first question is whether the attachment itself helps. To answer this, Table 1 compares, on three backbones, the full PhoneCLI system against its embedded pure VLM agent ( VLM-only ) — PhoneCLI with the compiled layer removed, all else identical. Table 1: Success rate, average steps, and tokens per task on AndroidLab. Model Condition Success Δ \Delta Avg. Steps Avg. Tokens Kimi-K3 VLM-only 68.8 % 68.8% + 0.8 +0.8 6.34 6.34 53.3 ​ k 53.3k PhoneCLI 69.6 % \mathbf{69.6%} 6.02 6.02 ↓5% 48.1 ​ k 48.1k ↓10% GLM-4.6V VLM-only 41.3 % 41.3% + 3.2 +3.2 7.07 7.07 61.4 ​ k 61.4k PhoneCLI 44.5 % \mathbf{44.5%} 6.52 6.52 ↓8% 54.4 ​ k 54.4k ↓11% Qwen3.7-Plus VLM-only 50.7 % 50.7% + 12.3 +12.3 7.52 7.52 40.1 ​ k 40.1k PhoneCLI 63.0 % \mathbf{63.0%} 6.72 6.72 ↓11% 34.4 ​ k 34.4k ↓14% Three observations follow: • (i) PhoneCLI adapts across backbones, with the largest gains on weaker ones. Attaching PhoneCLI improves the success rate by 12.3 12.3 points on Qwen3.7-Plus and by 0.8 0.8 points on Kimi-K3. Weaker backbones carry limited app-specific navigation knowledge, and the compiled map supplies exactly this missing knowledge. • (ii) Beyond the success-rate gains, tasks are also completed faster and more economically. On Qwen3.7-Plus, PhoneCLI reduces the average number of steps by 11 % 11% and the token consumption by 14 % 14% , as deterministic replay replaces multi-round visual navigation with sub-second, zero-token commands. • (iii) As backbone capability grows, the success-rate gain shrinks to 0.8 \mathbf{0.8} points on Kimi-K3, while the cost advantage persists — token consumption remains 10 % 10% lower and the step count remains 5 % 5% lower. An enhancement layer that keeps reducing cost even when it no longer adds capability is particularly valuable under the latency and cost constraints of mobile deployment. Table 2: Success rate (%) on AndroidLab. Method Size SR General-purpose LLMs GPT-4o – 31.2 Claude-Sonnet-4 – 40.6 GLM-4.6V 107B-A12B 41.3 Qwen3.7-Plus 35B 50.7 Gemini-2.5-Pro – 56.5 Kimi-K3 2.8T-A104B 68.8 GUI-specific models V-Droid 8B 38.4 UI-Tars-1.5 72B 38.4 UI-Genie-Agent 72B 41.2 MobileUse 72B 44.2 AutoGLM-Mobile 9B 46.8 AutoGLM-Phone 9B 47.7 AutoGLM-2024-10 – 36.2 PhoneCLI method w Qwen3.7-Plus PhoneCLI w/o Replay 35B 55.1 PhoneCLI 35B 63.0 Table 2 situates PhoneCLI against both general-purpose LLMs and GUI-specific models, and two points stand out. First, PhoneCLI lifts a modest backbone toward frontier level: Qwen3.7-Plus (35B) moves from 50.7 as a plain VLM agent to 63.0 with PhoneCLI, closing a 19.2 point gap against the 2.8T-A104B open-weight Kimi-K3 (68.8) to 5.8 points — and on the Kimi backbone itself, PhoneCLI still adds a further 0.8 points (69.6, Table 1 ). A 35B model enhanced by PhoneCLI is thus brought within reach of a 2.8T-class one, demonstrating the strong amplification PhoneCLI provides for weaker models. Notably, 63.0 is also above every GUI-specific model in Table 2 , even though those models are trained for GUI control while PhoneCLI is not. Second, the two PhoneCLI rows preview where the gain comes from. PhoneCLI w/o replay keeps the app map but removes the replay mechanism: the map is only injected as reference text, and the VLM must navigate screens on its own. Information injection alone lifts Qwen3.7-Plus from 50.7 to 55.1, while adding deterministic replay raises it further to 63.0 — the replay step contributes the larger share of the gain. This shows that PhoneCLI’s improvement is not merely from exposing map information to the VLM, but substantially from a dedicated design: executing that knowledge deterministically, without VLM involvement. Moreover, App. A.4 reports a generalization test in which our method is attached to the official AndroidWorld agent. 3.3 Ablation Studies We conduct ablations at two granularities: (I) how much each of PhoneCLI ’s two core components contributes—the compiled map used as information and deterministic replay—and (I) whether any individual mechanism within these components is dispensable. We study both questions on AndroidLab using Qwen3.7-Plus across nine apps. Figure 2: Success rate of the three conditions per app on AndroidLab (Qwen3.7-Plus), sorted by the information contribution (left: helpful, right: harmful). Inset: average steps and tokens on each condition’s own successful tasks. (I) Core decomposition. Three conditions decompose PhoneCLI’s stack along its modules. w/o map (VLM only) removes the compiled map together with its replay machinery, leaving Stage 3, the pure VLM agent, at 50.7%. w/o replay keeps the map but only as information, injecting its navigation reference into the VLM’s prompt, at 55.1% ( + 4.4 +4.4 ). PhoneCLI additionally executes matched commands deterministically, at 63.0% ( + 7.9 +7.9 ). Both components pay off, replay the more. Cost and coupling. Injection shortens navigation but leaves the VLM in the loop: every remaining round still re-reads the injected reference and still pays a full VLM call, so rounds fall while tokens stay flat (7.52 → \rightarrow 6.95; 40.1k → \rightarrow 39.6k). Replay removes the call itself — compiled navigation consumes no VLM call — and collapses a multi-step traversal into a single recorded operation (6.95 → \rightarrow 6.72; 34.4k, − 13 % -13% ). The rounds that remain are the actual interaction rather than navigation. All efficiency numbers average over each condition’s own successful tasks (Fig. 2 ). Failure modes. Figure 2 orders the nine apps by how much the injected information contributes, and the ordering is uneven: injection helps on some apps and hurts on others. Its failure mode is that every screen must pass through the model’s understanding, so it degrades exactly where the map is semantically sparse or noisy — which is what collapses Calendar to 0/14. Replay has no understanding step and recovers the same app to 4/14, but it executes the command as recorded, so a wrong landing is silent and only a post-replay check can catch it. Table 3: Ablation of PhoneCLI components on AndroidLab (Qwen3.7-Plus, 9 apps). Variant Success Δ \Delta Module 1: map annotation w/o element classification 47.1% − - 15.9 w/o semantic enrichment 47.8% − - 15.2 Module 2: routing and replay w/o landing check 50.7% − - 12.3 w/o landing hint 50.0% − - 13.0 w/o command completion 48.6% − - 14.4 PhoneCLI 63.0% – (II) Fine-grained mechanisms. Every variant in Table 3 removes one mechanism, and every removal costs at least 12 success-rate points. On the map side , element classification — the STABLE / DYNAMIC annotation the router uses to tell version-persistent elements from transient ones — is the costliest single removal ( − 15.9 % -15.9% ), and semantic enrichment, the aliases and page descriptions that make the structural graph routable to both the router and the verifier, costs nearly as much ( − 15.2 % -15.2% ). On the replay side , the landing check verifies that replay actually reached the target screen; without it, wrong landings silently mislead the VLM onto irrelevant screens ( − 12.3 % -12.3% ). The landing hint — the description and element list injected at handoff — keeps a teleported VLM from starting cold ( − 13.0 % -13.0% ). Command completion lets the OP path finish a replay-only task in one deterministic replay and one verification, with no VLM interaction at all; forcing such tasks through the VLM wastes that and adds risk ( − 14.4 % -14.4% ). Across The same ai evaluation question is explored in REFLEX with Jev for Efficient Selective..., which adds a research perspective. as detailed in the full paper on Arxiv The same ai evaluation question is explored in Jev-Mobile, which adds a research perspective.

Comments (0)

No comments yet

Be the first to share your thoughts!