SelfSearch: Reward-Free Search for Self-Improving Agents investigates whether an AI agent can improve its own implementation without repeatedly testing candidate versions on downstream tasks. The paper introduces a process in which agents study records of earlier self-modification attempts— including reasoning, tool use, code changes, and failures—then use that experience to revise their instructions, tools, and execution procedures. In the paper, the authors report that this reward-free search improves average benchmark performance across all six tested model–benchmark settings, while also reducing some execution costs.
What SelfSearch does
Many approaches to improving agent harnesses use downstream evaluation as a search signal. A candidate agent is run on development tasks, its score is measured, and the result guides the next revision. This can be expensive because every search step requires additional task execution. It also makes the search closely tied to the particular benchmark being used.
SelfSearch removes downstream task scores from the revision process. Instead, an agent modifies an editable copy of its own repository and leaves behind an episode record. The record includes its reasoning, tool actions, observations, code changes, and local verification results. Later agents read these records and use them as experience when deciding what to change.
The model weights remain fixed. What evolves is the surrounding agent implementation: its instructions, tools, orchestration, and code organization. The runtime that executes the model is kept outside the editable repository, so it continues to control inference settings, resource limits, and execution records.
This focus on accumulated experience distinguishes SelfSearch from approaches such as Meta-Skill for harness design — though the two lines of work differ in their use of evaluation feedback. SelfSearch learns from the process of self-modification itself rather than from downstream development-set scores.
How the search works
The search begins with an initial agent, called (B_0). Each self-improvement episode receives the current agent, read-only records from previous episodes, and a qualitative search direction. It then edits an isolated copy of the repository and verifies the changes using available tools and local checks. The edited copy becomes the next agent, while the episode record is saved for future revisions.
The experiments maintain two parallel lineages. The “capability” direction asks the agent to find limitations in its abilities, such as inefficient behavior, failed actions, or difficulty completing an operation, and then create reusable tools or procedures. The “adaptive” direction asks it to improve how it changes course when actions fail, evidence contradicts its assumptions, or a better strategy becomes apparent.
Both lineages share the records they produce. This means that an improvement discovered by one lineage can be inspected and reused by the other. For example, one agent may develop a trajectory-reading tool, while the other later copies and refines it after seeing evidence about its limitations.
The search runs for ten generations, with one episode per lineage in each generation. The resulting agents are evaluated only after the search checkpoints have been frozen. No downstream benchmark tasks or benchmark results are available to guide the revisions.
Results that stand out
The authors evaluate GPT and DeepSeek configurations on SWE-bench Verified, SWE-bench Multilingual, and Terminal-Bench 2.1. The population mean—the average success rate of the two final lineages—improves over the initial agent in all six model–benchmark combinations.
On Terminal-Bench 2.1, the GPT capability lineage rises from 43.8% to 55.1%, while the DeepSeek capability lineage rises from 65.2% to 73.0%. On SWE-bench Multilingual, the largest individual improvement is 6.7 percentage points with GPT and 5.0 points with DeepSeek. On SWE-bench Verified, both DeepSeek lineages increase success from 81.7% to 86.7%.
The improvements are not limited to accuracy. On Terminal-Bench 2.1, both DeepSeek lineages improve success while reducing average execution cost per task by 12.1% and 16.9%. On tasks solved by both the initial and evolved agents, the DeepSeek adaptive agent reduces execution cost by 38.5% on SWE-bench Multilingual while improving overall success by 5.0 percentage points.
SelfSearch also compares against linear and archive search baselines that use a ten-task SWE-bench development set. The reward-free method achieves comparable results with lower search costs. In the DeepSeek setting, generating both SelfSearch lineages costs $4.03, compared with $8.59 for linear search and $7.90 for archive search in the reported comparison. A harness produced by SelfSearch reaches 82.0% on Terminal-Bench 2.1 under the settings of a public nine-harness comparison, tying the reported top score.
The mechanism resembles the broader idea that agents can improve through structured accounts of their own experience, as explored by Retrospection-Only Fine-Tuning. However, SelfSearch changes the agent’s surrounding code and procedures rather than fine-tuning model parameters.
What the analysis suggests
Ablation results indicate that both ingredients matter: access to prior episode records and an evolving improver. Removing episode records lowers population-mean SWE-bench Verified success by 2.1 percentage points with GPT and 2.9 with DeepSeek. Keeping the initial agent as the improver throughout the search also reduces the mean by 1.7 and 2.9 points, respectively.
The traces show concrete forms of cumulative improvement. GPT agents develop file-viewing tools, bounded text search, structured trajectory readers, filters for locating relevant events, and links between tool calls and their results. DeepSeek agents repair a search tool that could hide matching text in the middle of long lines, then reproduce and verify the repair.
These findings support the authors’ hypothesis that self-modification exercises capabilities useful for downstream work: inspecting unfamiliar code, diagnosing failures, implementing changes, and checking results. They do not establish that reward-free search will improve every agent or task distribution. The experiments use two model configurations, three benchmarks, two lineages, and ten generations, and the final agents are evaluated only after search rather than selected by downstream performance. Within those reported settings, however, SelfSearch shows that an agent’s experience while improving itself can become a practical source of evidence for improving its later behavior.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!