EmbodiedRSI adapts the software around a robot foundation model while keeping that model frozen. Its central proposal is to choose the next experiment for the information it can provide about competing code and skill changes.
The EmbodiedRSI paper treats trials as a limited resource. Rather than repeatedly proposing similar repairs, the harness records hypotheses, selects experiments that distinguish them and carries useful experience into later adaptation. The authors report stronger results on two simulated benchmarks and a separate physical-arm evaluation.
Matched trials distinguish code changes from skill changes
The Fast System maintains a Hypothesis Graph with candidate code and skill hypotheses and their revisions. Each experiment states which hypotheses its outcome can test. Confidence changes with outcomes from matched initial states, while revision edges preserve how an explanation developed.
Experiment selection estimates the expected reduction in uncertainty relative to cost. The cost includes elapsed time, simulator steps, GPU use and resets. A high-value trial should clarify a relevant uncertainty without consuming disproportionate resources.
The system compares a code–skill combination with code-only and skill-only execution from the same starting conditions. It records a positive joint gain when the combination succeeds more often than either component alone. This gives the harness evidence for retaining a pairing rather than assuming two plausible changes will help together.
Memory learns from later adaptation
A Slow System retains raw execution experience, hypothesis evidence and reusable abstractions at different levels. It learns memory actions according to later task improvement, uncertainty reduction and experiment cost. The method includes actions such as merging, abstracting or skipping records.
The robot controller stays frozen while this external code, skill and memory layer changes. That boundary is important for interpreting the results: improved task execution does not mean the foundation model's weights learned new motor behavior during the evaluation.
On RoboCasa365, the authors report 77.0% overall success and 71.3% on Composite-Unseen tasks. The latter compares with 40.1% for Harness VLA, the strongest reported baseline on that suite. On LIBERO-Pro, overall success reaches 86.8%, compared with 72.1% for Harness VLA. Evolution and evaluation use disjoint seed sets, with ten trials per evaluation seed.
Physical transfer has a narrower meaning than training-free robotics
The real-world experiment deploys the simulation-evolved harness on an SO-101 arm with SmolVLA. Before evaluation, the authors collect 50 teleoperated demonstrations per task and fine-tune SmolVLA for 20,000 steps. The policy then remains frozen during trials. Zero-shot transfer therefore describes the harness transfer, not the absence of task-specific controller preparation.
Across five tasks with 30 trials per condition, overall success rises from 46.0% for SmolVLA alone to 71.3% with the harness. Performance does not improve on every task: glasses-bridge grasping falls from 60.0% to 56.7%.
These results concern the tested tasks and hardware. Choosing informative simulator trials and transferring external code can improve execution, but the paper's success rates do not establish deployment-wide physical safety or reliable performance on arbitrary household tasks. The failure cases and controller-preparation requirements remain part of the result.
Comments