Back to AI Research

AI Research

Mixture of Self-Improving Branches For Agent Harnes... | AI Research

Key Takeaways

  • Mixture of Self-Improving Branches For Agent Harness Optimization Agent performance depends not only on the underlying language model, but also on the softwa...
  • Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback.
  • Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy.
  • These constraints channel evolution along a single search trajectory, increasing the risk of converging to a local optimum.
  • We make the improvement process itself adaptive by organizing search into branches with evolving development subsets and proposal policies.
Paper AbstractExpand

Harness optimization provides a practical setting for recursive self-improvement (RSI), where agent-generated modifications inform subsequent changes through execution feedback. Recent work such as Meta-Harness implements this process through iterative code generation and evaluation, but retains a fixed development set and proposal policy. These constraints channel evolution along a single search trajectory, increasing the risk of converging to a local optimum. We make the improvement process itself adaptive by organizing search into branches with evolving development subsets and proposal policies. Each branch retains development cases solved by more of its leading harnesses than by those of other branches, drops cases solved by every leading harness across all branches, and revises its proposal policy using its own search history. To deploy the resulting complementary harnesses, we propose a router to select one development-selected branch head for each new input before execution. Across mathematical reasoning and agentic coding benchmarks, our system achieves relative improvements over Meta-Harness of 34.8% on Olympiad-level mathematical reasoning, 11.6% on Terminal-Bench 2.0, and 3.8% on SWE-bench Lite, with harness selection and router configuration based solely on development data. These results show that evolving branch objectives and proposal policies can yield complementary harnesses whose strengths a router combines without access to test outcomes.

Mixture of Self-Improving Branches For Agent Harness Optimization

Agent performance depends not only on the underlying language model, but also on the software “harness” that controls retrieval, tools, memory, feedback, and execution. This paper asks whether harnesses can improve more reliably if their search process itself evolves. Instead of optimizing one harness against one fixed development set, the authors create multiple search branches, give each branch a changing objective and proposal strategy, and then route new problems to the branch most likely to handle them.

What the paper changes

Earlier systems such as Meta-Harness repeatedly generate and evaluate new harness implementations on the same development cases, using a fixed proposal policy. This creates a single search trajectory. A candidate that performs well overall tends to guide future development, even if another candidate has a valuable capability on a smaller group of difficult cases.
The proposed method uses multiple branches that start from the same seed harnesses, development set, and initial guidance. Each branch then evolves independently. It maintains its own candidate history, execution traces, development subset, and proposal guidance. A branch’s objective is simply its average reward on the cases currently assigned to it, so changing the subset changes which harness behaviors are favored.
This builds on a broader line of work in which agents improve their execution environments without changing the underlying model. For example, Video-RSI’s harness evolution approach also uses execution feedback to revise an agent’s executable harness, but the present paper focuses on diversity across parallel search trajectories rather than on a video-understanding task. (Note: the supplied Franklin URL was not a valid external destination; therefore no link is included here.)

How branching and routing work

At regular intervals, the system examines the strongest candidates—or “frontier”—in every branch. For each development case, it counts how many frontier harnesses solve it. If one branch has a sufficiently large advantage over the others, that case remains in the winning branch’s subset and is removed from the others. The ownership margin is set to two in the reported experiments.
Cases that every frontier harness across all branches solves are removed from search. They no longer distinguish candidates or provide useful pressure for further improvement. Cases without a clear branch advantage retain their current memberships. Over time, this process gives branches different objectives based on their emerging strengths, without assigning predefined categories such as “geometry problems” or “coding failures.”
The branches also adapt how they propose new harnesses. Every five iterations in the default setup, the coding proposer reviews a branch’s recent search history and revises its guidance. The guidance records which modifications helped, which repeatedly failed, and which problems remain unresolved. As the development subsets diverge, these local lessons can produce different proposal priorities: one branch may continue refining a mechanism that helps its retained cases while another investigates different failure modes.
At deployment, the system does not simply choose the branch with the highest score. The branches have been evaluated on different subsets, so their scores are not directly comparable. Instead, the authors retain the best development-selected harness from each branch and train a router using development cases solved by exactly one branch head. For a new input, the router selects one head before that head executes, without seeing test outcomes or correctness feedback.

Results across reasoning and coding

The default experiments use two branches, 20 search iterations, a frontier of five harnesses, and a two-case ownership margin. The system is evaluated on Olympiad-level mathematical reasoning, Terminal-Bench 2.0, and SWE-bench Lite, with held-out test data used only for final evaluation.
On mathematical reasoning with Gemini 3 Flash, the routed system reaches 62.0% accuracy, compared with 46.0% for Meta-Harness. With Claude Sonnet 4.5, it reaches 30.5%, compared with 29.0%. On Terminal-Bench 2.0, task completion rises from 44.8% to 50.0%. On SWE-bench Lite, issue resolution rises from 63.6% to 66.0%.
These correspond to relative improvements over Meta-Harness of 34.8%, 5.2%, 11.6%, and 3.8%, respectively. The system also outperforms the strongest listed fixed baseline in each setting. The paper reports that the router performs comparably to the branch expert with the highest test accuracy and exceeds that expert in two settings, despite selecting experts using development evidence alone.
The reported results fit with other harness-optimization research, including Meta-Skill’s use of execution feedback to learn reusable principles for constructing harnesses. However, the present method’s distinctive mechanism is comparative branch coverage: development cases are reassigned according to which branch’s leading harnesses solve them, encouraging complementary search objectives.

What to keep in mind

The paper’s central claim is not that every branch becomes universally better. Rather, evolving objectives and proposal policies can produce complementary harnesses whose strengths are combined by routing. The ablations reportedly show the strongest performance when both development-subset updates and proposal adaptation are enabled, supporting the idea that the two mechanisms reinforce one another.
The selection procedure is deliberately development-driven, but the search still requires repeated execution of candidate harnesses and uses held-out test sets to measure final generalization. The branch heads are selected on different subsets, which is why direct score comparison is avoided. Router training also depends on development cases where exactly one head succeeds, so its effectiveness depends on the presence of sufficiently distinctive branch capabilities.
Finally, the experiments use two branches and specific update schedules, and the paper evaluates three benchmark settings rather than establishing universal gains across all agent tasks. The results nevertheless support a narrower conclusion: allowing the search process to specialize—and then selecting among specialized harnesses at input time—can improve deployable performance over a single fixed-objective evolution trajectory.

Comments (0)

No comments yet

Be the first to share your thoughts!