Back to AI Research

AI Research

Turbo Harness patches an agent's runtime for each task

Key Takeaways

  • Turbo Harness reuses earlier optimization traces to train a small editor that tailors an agent harness per instance.
  • The paper reports improvements across seven benchmarks under sp
  • The paper reports improvements across seven benchmarks under specific execution budgets.
  • An agent's runtime decides which tools it uses, what context it retains and how it checks its work.
  • Optimizing that runtime once can improve average performance, but different tasks may still benefit from different instructions or control logic.

An agent's runtime decides which tools it uses, what context it retains and how it checks its work. Optimizing that runtime once can improve average performance, but different tasks may still benefit from different instructions or control logic. Turbo Harness proposes adapting the runtime before each task begins.
In Turbo Harness: Instance-Adaptive Harness Optimization, Tunyu Zhang and colleagues reuse the records of a completed harness search. Those records become a playbook for a smaller model that proposes a task-specific patch. The main model then executes the task inside the revised harness.

Reusing the search archive

A global harness search produces more than its winning implementation. It also produces failed candidates, execution traces and evaluations that show which edits worked under particular conditions. Turbo Harness summarizes successful and unsuccessful strategies into a structured playbook instead of discarding that history.
The editor receives the task instance, the global harness source code and the playbook. It proposes code edits to the existing harness, rather than generating an entirely new runtime. The experiments use Qwen3.5-9B as that trainable editor.
Training applies several candidate patches for each training instance, runs the task with the frozen execution model, and uses the task results as reinforcement-learning rewards. Edits that cannot be applied or fail validation receive zero reward and are not executed. At inference, a failed patch application falls back to the global harness.

A single editing call before execution

The trained editor is called once per task instance. That makes the deployment procedure different from running a new multi-round harness search for every incoming request. It also separates changing the runtime from changing the model: the execution model's weights remain frozen.
The study covers seven benchmarks across interactive agent tasks, software engineering and terminal work. On SWE-smith-MR, the authors report pass rates rising from 50.7% to 64.0% with Claude Haiku 4.5, and from 70.7% to 88.0% with Gemini 3.7 Flash, compared with Meta-Harness. Those are gains of 13.3 and 17.3 percentage points.
The gains vary elsewhere. On the study's SWE-bench Verified split, Haiku rises from 56.7% to 59.3%, while Gemini rises from 38.4% to 54.4%. In the agentic experiments, Turbo Harness performs comparably to Harness-R1 on WebShop, rather than leading every baseline on every individual benchmark.

Why the execution budget matters

The coding evaluations impose a 40-step limit and a $3 cost limit per issue. The authors contrast that step budget with the 250 steps commonly used for SWE-bench Verified evaluation. Less room for repeated attempts makes runtime guidance particularly consequential.
Their Verified results cover a repository-stratified, held-out 150-issue test subset, not the entire benchmark's usual public evaluation. SWE-smith-MR uses 50 training and 50 test issues from 25 Python repositories. These details limit direct comparisons with scores obtained under other splits and budgets.
Turbo Harness shows how earlier optimization work can inform per-task decisions without retraining the executor. Whether it reduces total cost in a deployment still depends on the editor's overhead and how much execution work its patches save. Franklin AI has not independently reproduced the paper's results, and the reported improvements should remain tied to the tested models, tasks and budgets.

Comments (0)

No comments yet

Be the first to share your thoughts!