MedRSI: Recursive Self-Improvement for Medical Agents via Clinically Aligned Self-Evolution
Medical AI agents typically rely on a fixed set of tools designed by humans before deployment. Once these agents are in the field, they cannot autonomously learn from their mistakes or expand their diagnostic capabilities. MedRSI is a new framework that enables medical agents to perform "recursive self-improvement," allowing them to learn from their own diagnostic failures and autonomously develop new clinical tools. By incorporating principles from clinical practice, the system ensures that this evolution is safe, evidence-based, and focused on patient outcomes. The same ai evaluation question is explored in ScienceBuddy, which adds a research perspective.
How MedRSI Learns from Failure
The framework operates in a continuous cycle: the agent diagnoses a batch of patients, identifies its errors, and reflects on what capabilities were missing during those mistakes. Instead of just fixing errors, the agent uses this information to build new tools—such as specialized segmentation models or measurement code—that address the specific gaps in its knowledge. This allows the agent to evolve beyond its original design, discovering new ways to analyze medical data that its creators may not have initially anticipated.
Prioritizing Patient Safety
Because medical errors carry different levels of risk, MedRSI uses "clinical-cost-aware failure prioritization." Rather than treating all errors as equal, the system assigns higher priority to mistakes that could lead to serious patient harm. By focusing its limited development resources on these high-consequence failures, the agent ensures that its self-improvement efforts are directed toward the most critical clinical needs, rather than simply fixing the most frequent or minor mistakes. The same ai evaluation question is explored in MAPLE, which adds a research perspective.
Validating New Capabilities
To prevent the agent from adopting unstable or potentially harmful tools, MedRSI employs a "fast discovery with slow registration" process. When the agent invents a new tool, it is initially placed in an experimental pool. The tool is only integrated into the agent’s permanent toolkit after it has demonstrated consistent, sustained benefits across multiple subsequent patient cohorts. This approach mimics the clinical practice of validating new medical interventions before they are widely adopted, ensuring that the agent’s growth remains grounded in empirical evidence.
Performance and Results
In testing across public benchmarks for glaucoma and heart disease, as well as private clinical tasks, MedRSI outperformed manually engineered medical agents. The system successfully developed a wide range of capabilities, including advanced segmentation, measurement, and multimodal reasoning. These results demonstrate that medical agents do not need to remain static; through clinically grounded self-evolution, they can continuously accumulate new skills and improve their diagnostic accuracy based on real-world experience. The same ai evaluation question is explored in AutoViewMem, which adds a research perspective. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!