Back to AI Research

AI Research

ScholarEvolve uses research papers to propose and test agent runtime changes

Key Takeaways

  • ScholarEvolve organizes published methods into candidate improvements for agent harness modules.
  • Its reported benchmark gains test the current search procedure; continual improveme
  • Its reported benchmark gains test the current search procedure; continual improvement from future literature remains a proposed extension.
  • An agent that edits its own runtime can learn from failures, but those failures do not necessarily suggest the best new technique.
  • ScholarEvolve adds another source of proposals: published research that describes mechanisms the agent could incorporate into its tools, memory or execution workflow.

An agent that edits its own runtime can learn from failures, but those failures do not necessarily suggest the best new technique. ScholarEvolve adds another source of proposals: published research that describes mechanisms the agent could incorporate into its tools, memory or execution workflow.
Jingbo Yang and colleagues introduce the framework in Learning from Research: Toward Lifelong Agent Harness Evolution. The task model remains fixed while a research and coding process builds, combines and tests changes to the software around it.

Turning papers into bounded changes

The framework divides a harness into five modules: tool interfaces, context management, skills, memories and agentic workflows. That division gives each proposed change a defined responsibility, rather than allowing every candidate to rewrite the whole system.
Research-pool construction begins with execution trajectories. An auditor identifies recurring agent-attributable failures, and a research model turns them into capability gaps and literature queries. Topic modeling groups retrieved papers by mechanism so the candidate budget can cover different approaches instead of repeatedly selecting similar ideas.
For a chosen paper, a research agent reads the method and appendix, along with available official code, and produces a blueprint. A coding agent adapts that mechanism to the target module while keeping the other modules fixed. Health probes check imports, interfaces, artifact handling and valid actions before behavioral evaluation.
This is more than inserting a paper abstract into a prompt. The proposed method must become executable code that fits the host system's constraints.

Testing combinations before keeping them

A useful module change may interact badly with another one. ScholarEvolve first measures individual mutations, uses their estimated gains to shortlist combinations, then evaluates the assembled harnesses directly. Selection uses those observed gains, rather than assuming the separate improvements simply add together.
The method retains a new champion only when the lower endpoint of a paired bootstrap gain interval exceeds zero. It can also retain the original implementation of a module or keep the existing champion when no candidate demonstrates improvement.
Evolution, validation and final test sets are disjoint in the paper's formulation. Persistent skill libraries and memories are built from evolution trajectories and frozen for validation and testing, although modules can update local state within an episode. These boundaries matter when judging whether search has learned transferable behavior or memorized evaluation tasks.

Reported gains and the lifelong proposal

On AppWorld Challenge, the authors report Qwen3.5-27B task goal completion rising from 49.6% to 63.6%. On Tau2-Bench Telecom, GPT-5.4-mini pass@1 rises from 72.7% to 81.9%. The measurements use different tasks and metrics, so they should not be combined into a single general-purpose agent score.
AppWorld Challenge tests 417 tasks and includes Amazon or Gmail APIs absent from the evolution and validation sets. The Telecom test split contains 40 tasks involving tools and a simulated user. The paper reports means and standard deviations over three evaluation runs.
The lifelong component proposes refreshing the literature periodically and starting new search cycles from the retained champion. Current benchmark improvements do not establish that this process will keep improving indefinitely, or that every newly published technique benefits a deployed agent.
ScholarEvolve's contribution is a research-informed proposal process with implementation and validation gates. Applying a promising paper still requires adaptation, and the finished combination must earn its place through task-level evidence. Franklin AI has not independently reproduced the reported benchmark results.

Comments (0)

No comments yet

Be the first to share your thoughts!