Jev-Mobile is a new framework designed to make autonomous mobile agents more efficient. Currently, most mobile agents rely on powerful Vision-Language Models (VLMs) to make decisions at every single step of a task. This approach is often slow and expensive because it requires constant, heavy computation. Jev-Mobile changes this by splitting the workload: a VLM handles high-level planning, while a lightweight, specialized model called "Jev" handles the rapid, repetitive execution of actions on the screen.
How Jev-Mobile Works
The system operates through a delegation loop. The VLM reads the task, the current screen, and the history of actions taken so far. Instead of deciding every individual click, the VLM sets a "local goal" for the agent. Once this goal is set, the Jev model takes over. Jev uses the device's accessibility tree—a structured map of the UI elements—to identify available buttons and fields. It then executes a series of actions to achieve the goal without needing to consult the VLM again. After each action, the system refreshes its view of the screen to ensure the agent stays on track, and Jev continues until the goal is met or it hits a roadblock. The same computer vision question is explored in SlackDrive, which adds a research perspective.
Efficiency and Performance
By reducing the number of times the heavy VLM must be called, Jev-Mobile significantly improves performance. In tests on the AndroidWorld task suite, Jev-Mobile achieved a 79% success rate, which is competitive with other existing systems. More importantly, it demonstrated major gains in efficiency: for successful tasks, it reduced the average end-to-end execution time by 32.7% and cut the average model API costs by 73.4% compared to a standard step-wise VLM approach. The same large language models question is explored in CERA-MoA, which adds a research perspective.
Limitations and Future Outlook
While Jev-Mobile is effective, it has specific constraints. It relies heavily on the Android accessibility tree to identify interactive elements. If a target is visible on the screen but missing from the accessibility data, the agent cannot interact with it, as it currently lacks a visual fallback to "see" and click coordinates directly. Additionally, the system was tested specifically on Android mobile tasks, so its performance on web interfaces remains to be seen. Future research may explore replacing Jev with even smaller models or testing the framework across different types of digital environments. The same large language models question is explored in An Empirical Study of Harness Design..., which adds a research perspective. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!