Video-RSI: Recursive Self-Improvement of Video Understanding Agents via Harness Evolution
Video-RSI explores whether a video-understanding agent can improve its own way of gathering and using evidence—without changing the underlying model’s weights. The central idea is to let the same frozen language model play two roles: a Solver that answers video questions and an Editor that investigates failures, revises the agent’s executable “harness,” and tests whether the revision is worth keeping. In the paper introducing Video-RSI, the authors report that this process improves accuracy while often reducing the number of video frames processed.
What the paper is trying to improve
A video agent must decide what to observe before it can answer a question. Important events may be brief, far apart in time, or surrounded by visually similar content. The agent’s harness controls this process: it determines which video tools to call, how densely to sample frames, how to process observations, what to store in memory, and when to gather more evidence or answer.
A failure trace does not always explain what went wrong. If the agent misses an event, the problem might be that it never inspected the relevant moment, that it misinterpreted an observation, or that it saw the evidence but failed to use it during reasoning. These possibilities call for different fixes. Simply increasing frame sampling, for example, might recover missed events but make every future question more expensive.
Video-RSI therefore treats the harness itself as the object of improvement. The model weights remain fixed, while the executable code around the model evolves. This makes the approach related to work such as learning meta-skills for agent harness design with fixed model weights, but Video-RSI focuses specifically on diagnosing video-observation failures and retaining revisions based on both accuracy and visual cost.
How recursive harness evolution works
The process begins with a current harness running on training questions. The Solver produces answers, execution traces, and resource records. The trace includes tool calls, observations, and decisions, but it only shows what the current harness chose to observe.
The Editor then examines these outcomes and actively revisits the original training videos. It selects diagnostic queries designed to distinguish competing explanations. A query might inspect a different temporal region, use denser sampling, request a transcript or on-screen text, or ask a perceptual question about a particular observation. These additional observations are not used to answer new evaluation questions directly. They are gathered offline to understand why the existing harness failed.
Investigation continues as the Editor updates its diagnosis. If a denser inspection reveals that an event fell between sampled frames, the likely revision concerns temporal coverage. If the event was already visible, the Editor instead considers whether the observation was represented, remembered, or passed correctly into later reasoning. The goal is to turn individual failures into a reusable behavior change that can apply to future questions without relying on known answer locations.
The Editor converts the diagnosis into a candidate harness revision. Changes can affect prompts, tools, observation processing, memory, and execution control. The model can also run training-side checks before submitting the candidate for selection. Accepted changes become the starting point for the next attempt, creating a recursive loop in which improvements accumulate in code.
How candidates are selected
A candidate is not accepted merely because it fixes a few observed failures. Video-RSI evaluates the incumbent and candidate on a private, video-disjoint development set under matched conditions. The main metrics are multiple-choice accuracy and mean visual frame usage, defined as the average number of frames processed by successful visual-model calls during complete question-answering executions.
The selection rule permits two kinds of improvement. First, a candidate can be accepted when it increases accuracy while keeping cost growth within a fixed bound. Second, it can be accepted when it substantially reduces frame usage while allowing only a tightly bounded accuracy loss. In the reported setup, the authors use 20 revision attempts, allow up to 10% relative cost growth for an accuracy gain, and require at least a 20% cost reduction for the cost-saving branch.
This criterion reflects an important distinction between a locally helpful repair and a better overall agent. A more detailed tool may solve one failure but cause unnecessary visual processing on many other questions. Conversely, a change that improves evidence reuse may lower total frame usage even if it adds complexity inside the harness.
What results stand out—and what to keep in mind
The authors evaluate the final harness on four video-understanding benchmarks: MLVU, LongVideoBench, Video-MME, and EgoSchema. The harness is evolved using LVBench questions, with a separate private development set used for candidate selection. All model weights remain frozen; the Solver and Editor use DeepSeek-V4-Pro, while Qwen3.6-Plus provides visual observations.
Video-RSI reports accuracies of 72.9% on MLVU, 71.1% on LongVideoBench, 80.0% on Video-MME, and 78.2% on EgoSchema. Its corresponding mean frame counts are 61.1, 41.1, 21.6, and 42.4. Against the reproduced VideoSeek baseline, the reported comparison shows a 4.8-point accuracy gain and 27.2% fewer frames on MLVU. On LongVideoBench, accuracy is similar while frame usage is substantially lower; on EgoSchema, both accuracy and frame usage improve. Video-MME shows higher accuracy with a moderate increase in frames.
The ablations are intended to test two design choices: active investigation before editing and cost-aware candidate selection. The paper material reports that active investigation produces higher final accuracy than trace-only revision on MLVU at similar inference-time frame usage. The authors also compare against an accuracy-only gate to study the contribution of explicitly tracking visual cost.
These results establish that the reported harness-evolution process can improve the tested agent under the stated evaluation setup. They do not show that every video agent, model, or harness will improve in the same way. Evolution uses a particular model pair, benchmark mix, revision budget, and selection rule, and the final harness is trained through offline investigation on LVBench. The main conclusion is therefore specific: active diagnosis plus accuracy–cost selection is a promising way to improve video evidence acquisition without retraining the underlying models.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!