Back to AI Research

AI Research

Jev-Mobile: Jev as an Executor for Mobile GUI Agents | AI Research

Key Takeaways

  • Jev-Mobile is a new framework designed to make autonomous mobile agents more efficient.
  • This design allows multiple GUI actions to be executed under a single VLM decision, reducing expensive VLM inference while preserving adaptive interaction.
  • On the full AndroidWorld task suite, Jev-Mobile achieves 79% task success, compared with 78% for SeeAct-V and 84% for a Step-wise VLM baseline.
  • Among successful trajectories, it reduces mean end-to-end execution time by 32.7% and mean model API cost by 73.4% relative to Step-wise VLM.
  • These results show that decoupling high-level VLM reasoning from low-level action execution can substantially improve mobile GUI agent efficiency while maintaining competitive task performance.
Paper AbstractExpand

Vision-language models (VLMs) have become a common foundation for autonomous mobile GUI agents, but most existing systems rely on the VLM for both planning and action grounding at nearly every interaction step, leading to substantial latency and model-serving cost. We introduce Jev-Mobile, which shifts this paradigm to low-frequency VLM planning and high-frequency lightweight execution: the VLM specifies local goals, the accessibility tree defines a structured executable action space, and Jev, a fast typed decision model, repeatedly selects actions within this space. This design allows multiple GUI actions to be executed under a single VLM decision, reducing expensive VLM inference while preserving adaptive interaction. On the full AndroidWorld task suite, Jev-Mobile achieves 79% task success, compared with 78% for SeeAct-V and 84% for a Step-wise VLM baseline. Among successful trajectories, it reduces mean end-to-end execution time by 32.7% and mean model API cost by 73.4% relative to Step-wise VLM. These results show that decoupling high-level VLM reasoning from low-level action execution can substantially improve mobile GUI agent efficiency while maintaining competitive task performance.

Jev-Mobile is a new framework designed to make autonomous mobile agents more efficient. Currently, most mobile agents rely on powerful Vision-Language Models (VLMs) to make decisions at every single step of a task. This approach is often slow and expensive because it requires constant, heavy computation. Jev-Mobile changes this by splitting the workload: a VLM handles high-level planning, while a lightweight, specialized model called "Jev" handles the rapid, repetitive execution of actions on the screen.

How Jev-Mobile Works

The system operates through a delegation loop. The VLM reads the task, the current screen, and the history of actions taken so far. Instead of deciding every individual click, the VLM sets a "local goal" for the agent. Once this goal is set, the Jev model takes over. Jev uses the device's accessibility tree—a structured map of the UI elements—to identify available buttons and fields. It then executes a series of actions to achieve the goal without needing to consult the VLM again. After each action, the system refreshes its view of the screen to ensure the agent stays on track, and Jev continues until the goal is met or it hits a roadblock. The same computer vision question is explored in SlackDrive, which adds a research perspective.

Efficiency and Performance

By reducing the number of times the heavy VLM must be called, Jev-Mobile significantly improves performance. In tests on the AndroidWorld task suite, Jev-Mobile achieved a 79% success rate, which is competitive with other existing systems. More importantly, it demonstrated major gains in efficiency: for successful tasks, it reduced the average end-to-end execution time by 32.7% and cut the average model API costs by 73.4% compared to a standard step-wise VLM approach. The same large language models question is explored in CERA-MoA, which adds a research perspective.

Limitations and Future Outlook

While Jev-Mobile is effective, it has specific constraints. It relies heavily on the Android accessibility tree to identify interactive elements. If a target is visible on the screen but missing from the accessibility data, the agent cannot interact with it, as it currently lacks a visual fallback to "see" and click coordinates directly. Additionally, the system was tested specifically on Android mobile tasks, so its performance on web interfaces remains to be seen. Future research may explore replacing Jev with even smaller models or testing the framework across different types of digital environments. The same large language models question is explored in An Empirical Study of Harness Design..., which adds a research perspective. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!