Back to AI Research

AI Research

VISTA lets visual agents revisit original frames instead of relying on summaries

Key Takeaways

  • VISTA stores original visual observations and provides model-directed inspection tools.
  • Its public-game results and staged ablations show the value of revisiting evidence, with imp
  • Its public-game results and staged ablations show the value of revisiting evidence, with important differences between headline baseline and harness settings.
  • An agent may need a detail from an old screenshot only after several more actions reveal why it matters.
  • A textual summary can preserve the agent’s earlier interpretation while losing the visual evidence needed to revise it.

An agent may need a detail from an old screenshot only after several more actions reveal why it matters. A textual summary can preserve the agent’s earlier interpretation while losing the visual evidence needed to revise it.
VISTA addresses that problem with a visual harness around a multimodal model. It stores original observations outside the active conversation and lets the model retrieve frames, zoom into regions and inspect pixels as its reasoning develops. “Lossless” describes preservation of those returned frames, not a claim that the model’s internal representation retains every detail.

Keep the evidence available after context compaction

The harness stores every frame returned by the environment, including intermediate animation frames. Frames are indexed by turn and frame number and remain available independently of the current model context.
At each turn, the agent sees the current state and available actions. It can consult two notes: one for persistent understanding of the game and another for the current working state. Inspection tools can retrieve earlier frames, enlarge a selected rectangle or return pixel values. The model chooses when these extra observations are useful.
The harness itself is mechanical, not another trained model. It executes requested tools, returns their outputs and reports invalid requests. When the conversation approaches its context limit, the agent writes a handoff summary and resumes in fresh context; notes, stored observations and action history remain available.

What the public-game score means

The paper evaluates 25 public ARC-AGI-3 games, where agents discover unfamiliar rules and goals through interaction. Its Relative Human Action Efficiency metric compares completed-level action counts with those of first-time human players, weights later levels and gives unfinished levels no credit.
The authors report a score of 100.00 for VISTA with Claude Opus 5.0 and 99.00 with GPT-5.6 Sol. Both systems complete all 25 public games. For the former, the reported total is 7,302 actions versus the human reference’s 17,135. These are results for the paper’s public-game setup, not proof of unrestricted visual-world competence.
The headline official-baseline comparison also changes system settings. The Opus baseline uses high reasoning effort, whereas VISTA uses extra-high effort. The staged GPT ablations separately increase stopping limits, add a continuous conversation and notes, then add visual memory and inspection. The headline improvement should not be attributed entirely to one memory component.

Inspection adds more than a longer window

In the GPT ablation, adding stored visual memory and model-directed inspection raises the reported score from 70.05 to 94.10. Adding exact pixel readout raises it to 99.00. These steps are evaluated within the progressively assembled harness, rather than as independent guarantees for any agent.
Larger contexts are not automatically better in the reported tests. The default 200K-token setting scores higher than the larger tested setting while using fewer tokens. The paper also examines image scale and finds that simply supplying more image tokens need not improve performance.
The authors apply the harness to browser-game benchmarks and a static visual-tracking benchmark as well. Those tests broaden the evaluation, but they do not turn a public-game result into evidence of reliable general computer use.

A related question: when to gather new evidence

WorldAuditBench studies agents investigating defects in interactive 3D scenes. It distinguishes adaptive investigation from analyzing a fixed recorded trajectory. VISTA instead focuses on preserving and revisiting observations while solving a task. Together, the studies make evidence access a concrete design question: can the agent obtain the observation its next inference needs, or is it limited to what an earlier pass happened to preserve?

Comments (0)

No comments yet

Be the first to share your thoughts!