Franklin AI Explainer

Google Open-Sources RRSI for Self-Improving AI Agents

Key Takeaways

  • Shows how AI agents can improve prompts, tools, memory and workflows without retraining model weights.
  • Adds leakage, noise and cost controls aimed at making harness evolution transfer beyond benchmark tasks.
  • Gives developers an Apache 2.0 research framework for testing self-improving agent systems.

Google Cloud AI Research, in collaboration with researchers from UNC-Chapel Hill, Stanford and Washington University in St. Louis, has released RRSI, a framework that allows AI agents to improve their operating harness without changing the underlying model weights. The Apache 2.0 project can modify prompts, tools, memory, control flow and sub-agents while using regularization to reduce benchmark overfitting, according to MarkTechPost’s report.
Formally called Regularized Recursive Self-Improvement, RRSI is available on GitHub and supports any LiteLLM model string. Its default configuration assumes Claude Opus 4.8 running through Vertex AI. The project is presented as a research framework rather than an official Google product.

RRSI targets a weakness in agent evolution

Many harness-evolution systems repeatedly propose changes, evaluate them on a fixed set of tasks and retain the highest-scoring version. Reusing the same tasks can cause a harness to memorize benchmark-specific details instead of learning general strategies.
The RRSI paper identifies three related failure modes: fitting to a particular benchmark, chasing evaluation noise and accumulating unnecessary complexity. These problems can produce large gains on the evolve set while reducing performance on tasks the system never used during optimization.
RRSI addresses the search process rather than changing the model itself. Every major harness component remains editable, but candidate changes must pass controls designed to test whether an improvement is meaningful, affordable and transferable.

How the framework controls self-improvement

On the proposal side, RRSI uses an annealed edit budget. Early rounds can combine several edits to explore broader mechanisms, while later rounds narrow the process to a single attributable change. The schedule follows a cosine curve, moving from larger edits toward smaller ones over the course of a run.
The system also maintains an evidence ledger for each candidate. The ledger records the edited component, the hypothesis behind the change, score movement and cost movement. The proposer can consult this history, allowing the process to avoid repeatedly pursuing ideas that previous evaluations have already weakened. When progress stalls within the measured noise range, RRSI shifts exploration toward components that have not yet been modified.
Selection adds further safeguards. A leakage critic screens proposed diffs for task names, entities, answers or benchmark-specific logic before evaluation. A noise-adjusted floor uses repeated runs of the unchanged base harness to estimate expected variation. A candidate must clear that uncertainty rather than relying on a small apparent gain.
RRSI also applies a cost rule: additional inference tokens must be justified by measured performance improvements. Components that stop producing gains become candidates for pruning. The researchers compare these controls with familiar regularization ideas, mapping the edit budget to L0 regularization, pruning to Lasso or L1, and the cost rule to Ridge or L2.

Reported results include held-out improvements

The reported results show gains on both evolve sets and evaluations that were not used for selection. On Terminal-Bench 2.1, the system improved from 74.2% to 80.2%. On SWE-bench Verified, which the source says was never used for selection, performance increased from 82.0% to 83.8%.
RRSI also improved results on several out-of-distribution benchmarks: JobBench rose by 4.7 points, GDPval by 3.5 points and APEX-Agents by 3.7 points. EngDesign improved by 4.9 points on its evolve evaluation, while Frontier-Eng gained 4.3 Medal points. Harvey LAB increased by 1.1 points on the evolve split and 2.3 points on its held-out split.
All six held-out splits in the reported evaluation improved. Using Gemini 3.5 Flash as the policy model, Terminal-Bench 2.1 rose from 64.6% to 78.7%, while SWE-bench Verified moved from 76.8% to 79.0%.
The framework also used fewer policy tokens per trial on the agentic workspace instance: 2.42 million compared with 3.80 million for unregularized evolution. The source describes that reduction as 30% in the paper’s abstract and 36% on the project page, so the reported percentage depends on which project description is used.

RRSI trades peak evolve scores for transfer

In the comparison reported by the RRSI team, Meta-Harness achieved the highest Harvey LAB evolve score at 93.0, compared with 90.5 for RRSI. However, RRSI recorded the strongest out-of-distribution average among the methods listed: 43.6, versus a starting-harness baseline of 39.7. Meta-Harness reached 40.6, while the other systems remained at or below the baseline.
That result reflects the project’s central design choice. RRSI does not simply retain the most successful candidate on the optimization split; it also penalizes leakage, unexplained noise, increased cost and ineffective complexity. The tradeoff is that a more conservative system may give up some peak evolve-set performance in exchange for stronger transfer.
The framework can be installed with Python 3.10 or newer. Its coding instance additionally requires Docker and Harbor, and new domains connect through a single adapter module. Each round creates two candidates in separate Git worktrees, screens and evaluates them, then advances the branch to the selected winner.
RRSI’s open-source release gives researchers a way to test recursive harness improvement while keeping model weights fixed. Its reported held-out gains are promising, but the project remains research-grade, and the evaluation settings, cost rules and regularization parameters will determine how well the approach transfers to other agent domains.

Our read

Franklin AI Take

RRSI is notable because it treats agent improvement as a software-search and evaluation problem rather than a model-training problem. Its key contribution is not merely allowing agents to edit their harnesses, but constraining those edits with evidence tracking, noise estimates, leakage checks and cost controls. The reported held-out gains are encouraging, while the differing token-reduction percentages show why readers should inspect the evaluation details. Results will likely depend heavily on benchmark design, regularization settings and the adapters used for each domain.