What the paper is about
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.
What it covers
Shockingly Simple Self-retrospection Improves Agentic Models Without RL Jonathan Light Christopher Zhang Cui Jeonghye Kim Roger Creus Castanyer Emiliano Penaloza Zhengyan Shi Alessandro Sordoni Marc-Alexandre Côté Xingdi Yuan Minseon Kim RPI Affiliation: UC San Diego Affiliation: KAIST Affiliation: Mila Affiliation: Microsoft Research 🖂 Corresponding author: [email protected] September 2026 Abstract People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone . The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO’s 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer. Figure 1: ROFT achieves competitive performance with 63% less training time than GRPO . Held-out performance of ROFT versus GRPO on agentic coding tasks. 1 Introduction We sometimes understand an experience differently in the act of recounting it. Explaining why a conversation went badly, someone might begin with “I did not make my point clearly enough.” But as they reconstruct the exchange, another explanation emerges: they kept defending their proposal while the other person was questioning its premise. The lesson is no longer to explain the same point more forcefully, but to establish which question needs answering. Nothing about the original outcome has changed. What changes is their understanding of what happened—and, with it, how they might act next time. For language-model agents, this suggests a complementary training target: not only the actions taken during an attempt, but retrospective explanations that make sense of the experience. Reinforcement learning with verifiable rewards (RLVR) offers an effective way to learn from experience by reinforcing task-solving behavior according to its outcomes ( Lambert et al., 2024 ; DeepSeek-AI et al., 2025 ) . Methods such as GRPO use relative rewards among sampled attempts to update the policy ( Shao et al., 2024 ) . Yet a single outcome reward, success or failure, does not identify which assumption was mistaken, which decision mattered, or when a correction should apply. This limitation is especially relevant when rewards are sparse: an all-failure group has no within-group binary-reward contrast, even though its trajectories may contain useful evidence. Natural language can express an interpretation of that evidence, connecting observations to decisions and stating lessons that may apply elsewhere. Prior work has used reflections in several ways: as context for later attempts ( Shinn et al., 2023 ) , or to obtain successful reasoning and reflection-informed solutions for training ( Zelikman et al., 2022 ; Shi et al., 2026 ) . Critique fine-tuning, which trains on critiques as a form of reflection, demonstrates that teacher critiques can be transferred offline to a student to improve its question-answering performance ( Wang et al., 2025b ; Wang et al., 2026 ) . However, it remains unclear whether training only on self-generated reflection can improve agents. The central question is therefore: Can an agent improve its future actions by training only on self-generated retrospective explanations of its own experience? We call this direction Retrospection Reinforcement (RR) : training on self-generated retrospections of past experience, with the aim of improving subsequent behavior. A retrospection may summarize events, track how beliefs evolved, explain decisions or outcomes, identify corrections, or articulate lessons for future attempts. Our procedure trains on this retrospective text rather than directly supervising the recorded task-solving actions. The hypothesis is that learning to explain an experience can improve how the model acts in later attempts. We assess this hypothesis through subsequent task performance and behavior changes, not through the fluency of explanations or a reduction in their prediction loss. To study this route in isolation, we introduce Retrospection-Only Fine-Tuning (ROFT) , a deliberately minimal online procedure. The agent attempts a task, observes the outcome and available feedback, and generates a retrospection of its own attempt. We then fine-tune the same model with a next-token prediction loss on the retrospection tokens only. The task, attempted actions, and observations serve as context, not prediction targets. Both successful and unsuccessful attempts are eligible, and the procedure uses neither an external teacher nor a reward-based policy update. Subsequent attempts receive no stored retrospection and require no additional reflection step: any benefit must transfer through the updated weights. These exclusions make simplicity an experimental choice, isolating what retrospection training can contribute without direct action supervision. Our software-engineering experiments provide evidence that learning to explain can improve learning to do. First, when trained on problems with both successful and unsuccessful base-model attempts—the learning zone —ROFT achieves higher held-out solve rates than GRPO at the reported checkpoints, with faster early progress in wall-clock time and sampled solution attempts ( ), without a verifier. Second, it learns to solve individual tasks on which all 64 sampled base-model attempts failed, extending learning beyond the frontier without an initially successful trajectory or variation in binary outcomes ( ). Third, behavioral analyses find changes consistent with indirect credit assignment: likelihood increases are more selective for correct than incorrect recorded turns. Changing the retrospection prompt to emphasize more direct solutions also produces shorter subsequent attempts, despite the absence of an explicit length penalty or action-target loss ( ). These results suggest a complementary design axis for agent training: not only which experiences to learn from, but what to learn to say about them. Our results show that experience can support useful training targets even when binary outcomes provide no contrast or in the absence of any outcome verdict. The prompt-dependent behavioral changes further suggest that the content of those targets matters: changing what an agent emphasizes in retrospection can change how it subsequently acts. Retrospection is therefore not merely a record of experience, but a potentially steerable source of supervision. 2 Retrospection-Only Training An unsuccessful attempt may be a poor example to imitate but still provide useful experience to interpret. Here, we describe how an agent can be trained exclusively on retrospections of its own experience and evaluated on its subsequent task performance. 2.1 Learning to explain, evaluated by doing Learning to do directly optimizes task-solving outputs, including reasoning, answers, and actions. Reward-based updates, imitation, and distillation can all serve this purpose. Learning to explain instead trains the model to produce a retrospective account of a completed attempt. We call this direction Retrospection Reinforcement (RR) . A retrospection is a textual account of past experience, generated in light of the recorded interaction and any available feedback. Its form can range from a trajectory summary or a ledger of evolving beliefs to an explanation of decisions or outcomes, a correction, or a reusable lesson. No particular content structure is required. Retrospections may discuss actions and possible corrections; the distinction is between training on a post-attempt account and directly supervising task-solving actions. contrasts direct action training with retrospection prediction, where the recorded attempt supplies context rather than prediction targets. Our hypothesis is that learning to explain an experience changes the knowledge and representations available when the same model next acts. Below, we describe a minimal algorithm for studying this hypothesis. Figure 2: Illustrative RL versus RR. Black arrows show RR’s online collection and rendering, not RL updates. Orange marks targets; yellow, masked context. 2.2 ROFT: A minimal test of explanation-to-action transfer Retrospection-Only Fine-Tuning (ROFT) repeats a simple loop: attempt a task, explain the experience, fit the explanation, and act again using the updated weights. We exclude direct action supervision and retained retrospection so that neither can account for an improvement in doing. Collect experience. The agent attempts a task x x , producing an interaction trajectory τ \tau , and receives available feedback v v . In our software-engineering experiments, the record includes actions and observations, the submitted patch, and available test feedback. Both successful and unsuccessful attempts are eligible: the method does not require a successful example to imitate. Explain the experience. The same model samples K K retrospections from a context h ( x , τ , v ) h(x,\tau,v) containing an instruction and a bounded rendering of the record. Our default experimental prompt asks for a consequential assumption or decision, supporting or contradicting evidence, a correction when appropriate, and concrete triggers for applying the lesson. This is one instantiation of retrospection, not a requirement on its form. The model supplies its own interpretations rather than receiving an external teacher’s critiques. Available verdicts condition these interpretations; they do not become reward weights in the objective. The example below shows how even a passing attempt can yield a lesson about a specific decision. Retrospection after a passing attempt (full prompts and responses in ) System prompt (abridged). Identify a consequential assumption or decision, the evidence for or against it, a correction if needed, and concrete triggers for applying the lesson. State uncertainty when evidence is insufficient. User prompt (abridged). Task: Prevent a failed Pact context-manager test from leaving stale interactions that cause later tests to fail. Action: Inspect the context manager’s exit and verify() methods. Observation: verify() clears the interaction list, but exit skips it when an exception occurs. Action: Clear self.interactions on exceptional exit, then run the Pact and consumer tests. Observation: All 70 tests pass. Retrospection (final-answer excerpt). “The correct decision was to clear interactions directly in exit when an exception is detected, ensuring proper state reset regardless of whether verify() is called.” Fit only the explanation, not the attempt. Let 𝒟 t \mathcal{D}{t} contain the retained context–retrospection pairs ( h , y ) (h,y) for update t t . Holding these generated targets fixed, we fine-tune the model π θ \pi_{\theta} with next-token cross-entropy: ℒ ROFT ( θ ; 𝒟 t ) = − 1 T t ∑ ( h , y ) ∈ 𝒟 t ∑ j = 1 log π θ ( y j ∣ h , y < j ) , T t = ∑ ( h , y ) ∈ 𝒟 t | y | . \mathcal{L}{\mathrm{ROFT{}}}(\theta;\mathcal{D}{t})=-\frac{1}{T_{t}}\sum_{(h,y)\in\mathcal{D}{t}}\sum{j=1}\log\pi_{\theta}(y_{j}\mid h,y_{<j}),\qquad T_{t}=\sum_{(h,y)\in\mathcal{D}_{t}}|y|. (1) The loss averages globally over retained retrospection tokens. The task, attempted actions, observations, and feedback are masked as prediction targets, although gradients can flow through their context representations. There is no action-target loss, reward-based policy update, or independent semantic-quality filter on the retrospections. Act again using updated weights. The learner collects fresh attempts and retrospections as training proceeds. Subsequent attempts, including evaluation, receive neither stored retrospection text nor an added retrospection step. Any benefit must therefore transfer through the updated weights, rather than through access to a written lesson. This closes the loop between explaining and doing while keeping their training targets distinct. Sampling settings, target processing, batch construction, and asynchronous weight refresh are specified in . 3 Results We evaluate explanation-to-action transfer through three questions. Does retrospection-only training improve held-out solve rates, and at what training cost? Can learning begin when all sampled base model attempts fail? Does retrospection content shape subsequent behavior as measured by action-likelihood changes, interventions on retrospection instructions, and changes to content weighting? These questions assess behavioral transfer. We distinguish two learning regimes using sampled base-model outcomes. A problem is in the learning zone if the base model produces both successful and unsuccessful solutions across k k sampled attempts; it is beyond frontier if none of those attempts succeeds. illustrates the distinction for Qwen3.5-4B on all 500 SWE-bench Verified problems using 16 attempts per problem. The experiments below use three attempts to screen the learning-zone training subset and 64 attempts to establish the all-failure starting points for single-task training. Figure 3: SWE-bench Verified problems categorized by base model outcomes over 16 attempts. Percentages of problems solved in all, some, or none of 16 attempts pooled from 16 evaluations. 3.1 Learning Zone We first evaluate generalization and efficiency in the learning zone by comparing ROFT and GRPO initialized from Qwen3.5-4B and trained on SWE-rebench-767, a subset of 767 problems from SWE-rebench ( Badertdinov et al., 2025 ) . Our training subset is specifically selected to favor GRPO : because its reward-derived policy-gradient signal requires variation in rewards within a group, we retain only problems on which Qwen3.5-4B produces both successful and unsuccessful solutions across three initial attempts. We assess held-out performance on SWE-bench Verified ( OpenAI, 2024 ) and SWE-bench Pro ( Deng et al., 2025 ) . All three datasets require agents to resolve real-world repository issues by producing code patches evaluated against executable tests. SWE-bench Verified comprises 500 human-validated Python tasks, whereas SWE-bench Pro emphasizes more complex, long-horizon tasks that often require substantial changes across multiple files. Qwen3.5-4B performance is 44.2% and 23.9% on Verified and Pro respectively in our setup. Generalization. ROFT achieves higher held-out solve rates on both benchmarks in : its 20-update checkpoint reaches 49.2% on SWE-bench Verified and 29.0% on SWE-bench Pro, compared with 48.0% and 25.3% for GRPO at 40 updates. These gains do not require continued improvement in training reward. ROFT’s training curve levels off around 20 updates, whereas GRPO’s reward continues to improve with longer training. Yet GRPO’s Verified solve rate falls from 48.0% at update 40 to 46.0% at update 90 ( ), a pattern consistent with overfitting to the training tasks. We hypothesize that ROFT’s early saturation reflects increasingly repetitive retrospections as training revisits familiar tasks, reducing the new supervision they provide, which is consistent with decreasing reflection entropy and SFT loss. See for details. Training efficiency. Starting from Qwen3.5-4B and training on SWE-rebench-767, ROFT makes faster early progress than GRPO in terms of wall-clock time, optimizer updates, and sampled solution attempts ( ). With the same allocation of four training and four inference GPUs, ROFT completes 40 updates in 4.32 hours using 6,035 solution attempts, compared with 8.30 hours and 11,576 attempts for GRPO, i.e. approximately half the training time and half as many solution attempts. This early advantage also transfers to held-out performance: ROFT solves 49.0% of SWE-bench Verified after ten updates and 1.29 hours, exceeding the 48.0% achieved by GRPO’s 40-update checkpoint, which takes 8.30 hours ( ). Thus, ROFT attains a higher observed test score in roughly one-sixth the training time. Comparison with prior methods. We also compare ROFT with baselines from prior work, with results in and method details in . ROFT can also be effectively combined with GRPO by adding a retrospection step to each GRPO rollout ( ). (a) Training reward and held-out evaluations. (b) Held-out performance. Figure 4: Early training gains transfer to held-out tasks, while later training reward need not improve generalization. (a) Training reward through 40 ROFT updates and 100 GRPO updates. Flags report separate SWE-bench Verified test set solve rates at the marked checkpoints. (b) Solve rates on SWE-bench Verified and SWE-bench Pro for ROFT at update 20 and GRPO at update 40. (a) Training wall time. (b) Optimizer updates. (c) Sampled solution attempts. Figure 5: ROFT has higher training efficiency. ROFT and GRPO trained on SWE-rebench, with training reward plotted against (a) wall time, (b) optimizer updates, and (c) completed solution attempts, including zero-advantage attempts rejected by GRPO. See for details. 3.2 Learning Beyond the Model’s Frontier We next ask whether learning can begin without any observed successful base-model attempts, rather than only on the mixed-success problems. With binary success rewards, an all-failure GRPO group has zero within-group advantage and provides no reward-derived policy-gradient signal, so GRPO does not learn in this regime. ROFT instead obtains supervision from retrospections of failed attempts. To study an all-failure starting point, we train separate Qwen3.5-4B runs on individual SWE-bench Verified and SWE-rebench tasks, each beginning with 0/64 successful base model attempts on that task and using only self-generated retrospections for supervision. After 40 updates, the online solve rate reaches 1.75% on SymPy and 3.29% on SymbiFlow ( ). Thus, retrospection-only training can produce successful solutions from an all-failure starting point. Training on a single problem with ROFT does not degrade test-set performance significantly ( ). (a) SymPy 20916 (SWE-bench Verified). (b) SymbiFlow 17 (SWE-rebench). (c) Pylint 6386 (SWE-bench Verified). Figure 6: ROFT can improve performance on near and beyond frontier problems. Online solve rates during 40 updates of ROFT on a single task. SymPy and SymbiFlow begin with 0/64 successes; Pylint begins with 2/64 and reaches a smoothed solve rate of 12.00% at update 40. We see similar results on Django and with longer training runs ( ). 3.3 Behavioral Analysis Credit assignment. Can retrospection-only training assign credit to individual actions? We evaluate a fixed set of base model trajectories on SWE-rebench-767, with assistant turns labeled for correctness by GPT-6 Astra using the recorded evidence. For each turn, including both reasoning and actions, we compare its mean token log-likelihood under the initial Qwen3.5-4B model and after ten optimizer updates of ROFT or GRPO, conditioning on the same recorded history. shows the fraction of turns whose likelihood increases, computed separately for each correctness label within each trajectory and then averaged equally across trajectories containing that label. Under RR, correct turns increase in likelihood more often than incorrect turns (32.41% versus 21.21%), whereas GRPO shows similar rates for both (44.14% versus 45.30%). This separation is consistent with indirect credit assignment without direct action supervision. See for details. (a) Turn likelihood increases. (b) Rollout lengths at update 10. (c) ROFT combined with GRPO. Figure 7: (a) RR conducts more credit assignment than GRPO. Bars show trajectory-balanced fractions of recorded turns whose mean token log-likelihood increases relative to the initial model after ten optimizer updates. Under RR, correct turns increase in likelihood more often than incorrect turns; GRPO shows little separation. This pattern is consistent with indirect credit assignment from retrospection-only training. We attribute the lower likelihood-increase fractions under RR to its larger policy shift over the same ten updates. (b) We can induce behaviors such as completing the task faster through RR alone without any length penalties. Mean tokens and turns per trajectory under the baseline ROFT prompt and a variant that prompts the model to retrospect on how to solve the task faster. (c) ROFT can be combined with GRPO. Update 20 performance. See and . Inducing behavioral changes through retrospection. Can the content of retrospection steer subsequent task-solving behavior? We compare standard ROFT with a variant that asks the model to identify how it could have solved the task more directly after a successful attempt while preserving correctness. Both runs use the same initialization, training data, solver prompt, and retrospection-only objective. The variant uses 11.7% fewer generated assistant tokens and 13.1% fewer assistant turns than the baseline ( ). These observations are consistent with content-directed behavioral change from training on retrospections alone, without an explicit length penalty or action-target supervision. gives the experimental procedure, complete prompts, and detailed results. (a) KL conditioned on role. (b) KL conditioned on tool. (c) Retrospection impact. Figure 8: (a, b) Reasoning and read/search contexts show larger prediction changes. KL divergence on fixed trajectories before and after one ROFT update, grouped by token roles and tools. (c) Evidence, correction, and attention weighting yield the largest observed solve-rate gains. Solve-rate changes relative to uniform weighting (Uniform) after one update with doubled category weights or continuous prompt-attention weights (Attention-up), normalized to mean one. Points show means with ± 1 \pm 1 bootstrap SE; evaluation uses ten repetitions on each of the 64 training problems (640 attempts per model). Details in . Fine-grained analysis of action-trajectory changes. Where does learning to explain change the model’s subsequent action predictions? We isolate a single retrospection-only update of Qwen3.5-4B on 255 retrospections from 64 source attempts, then compare the full next-token distributions before and after training on the same original rollout prefixes. Among the mechanically identified roles, reasoning has the largest task-balanced mean forward KL, a 3.45 × 3.45\times difference from tool calls ( ). Changes are also concentrated earlier in the rollout: the first normalized assistant-turn decile has the largest mean KL ( ). Read/search contexts shift more than modification contexts: the task-balanced mean KL for Glob and Read is 1.35 1.35 – 1.89 1.89 times that for Write and Edit across the four pairwise comparisons. These patterns are consistent with transfer to reasoning and information gathering despite the absence of action-target supervision. See for details. Weighting retrospection contents. Does emphasizing particular retrospection content improve subsequent task solving? Using the same frozen retrospections, we compare uniform target weights (Uniform) with variants that double the relative weight of evidence, task-specific corrections, reusable lessons, or thinking tokens. GPT-6 Astra labels the evidence, correction, and lesson categories; thinking tokens are identified mechanically. Attention-up instead weights each nonstructural target token by one plus the base model’s mean attention mass to the user prompt when predicting it. All variants receive one update from the same base model, with weights normalized to mean one across the target batch. We evaluate each checkpoint on the same 64 training problems with ten paired evaluation seeds, yielding 640 attempts per model. Attention-up, correction-up, and evidence-up have the highest observed solve rates among these variants: 54.69%, 54.38%, and 54.22%, respectively, versus 50.63% for Uniform ( ). Why might retrospection training work? We hypothesize that RR improves behavior by refining the model’s understanding of an experience rather than directly reinforcing individual actions. Revising an assumption or learning when a correction applies may influence many future decisions, providing a route to generalization and potentially explaining the faster early progress relative to GRPO. Such higher-level supervision can also be more fine-grained than a trajectory-level reward: evidence grounds an explanation in observed behavior, while corrections identify which decisions should change and why. Correct and incorrect turns show a larger gap in likelihood-increase rates under RR than under GRPO, consistent with more selective credit assignment ( ). Mechanistically, predicting retrospection tokens can backpropagate gradients through attention to representations of the recorded actions and observations, updating shared model parameters even though the trajectory tokens themselves carry no prediction loss. This provides a pathway for explanation training to change subsequent actions without directly supervising them. Consistent with this, giving greater weight to retrospection tokens whose predictions attend more strongly to the user prompt containing the trajectory yields a higher observed solve rate than uniform weighting ( ). These findings motivate future work to trace how retrospection gradients reshape action-relevant representations and to test the role of this pathway in learning speed and generalization. 3.4 Ablating the Retrospection Learning Signal Number of retrospections per rollout. Four retrospections per rollout give the highest observed test performance in this sweep, suggesting a useful balance between rollout diversity and retrospection diversity ( ). We compare 2, 4, and 8 independently sampled retrospections per rollout, using 128, 64, and 32 source rollouts, respectively, to keep 256 nominal retrospection targets per update. Fewer retrospections expose training to more source trajectories, whereas more retrospections provide more independently sampled interpretations of each trajectory. Neither extreme improves on the four-retrospection control here. See for details. (a) Online, off-policy, and offline. (b) Retrospections per rollout. (c) Qwen3.5-9B at 20 updates. Figure 9: (a) Online RR has the highest observed solve rate. SWE-bench Verified performance from Qwen3.5-4B with online, off-policy, or offline retrospection training. Off-policy reflections freeze the retrospection generator; offline RR freezes both the attempt and retrospection generators before training. (b) Four retrospections per rollout perform best in this sweep. Solve rates when varying retrospections per rollout and inversely varying source rollouts, holding the nominal retrospection batch fixed. (c) The advantage over GRPO also holds for Qwen3.5-9B. SWE-bench Verified solve rates under ROFT and GRPO from the same larger base model. Details in . Different model. The advantage over GRPO also holds for Qwen3.5-9B: after 20 updates, ROFT solves 58.8% of SWE-bench Verified versus 55.8% for GRPO ( ), using 1.93 versus 5.33 training hours. For Qwen3.5-9B, training performance begins to plateau around 20 updates. gives the training and evaluation protocol. Verdict vs. no verdict. Providing the final correctness verdict does not improve the observed SWE-bench Verified score in this comparison: after 20 updates, both the verdict-conditioned baseline and the no-verdict variant achieve 49.2% ( ). The no-verdict variant generates retrospections from the trajectory summary and final patch, without the final grader feedback. We hypothesize that environmental feedback already present in the trajectory, including observations from the agent’s own tests, provides sufficient context to generate useful retrospections without an explicit final correctness label. describes the training and evaluation protocols. Off-policy reflections and offline. Does it matter whether the model learns from its own current retrospections? compares three training regimes at 20 updates. Online RR, in which the evolving learner generates both attempts and retrospections, achieves 49.2%. Freezing the retrospection generator at the base model while continuing to refresh learner attempts yields 47.4%. Fully offline RR instead trains on a fixed corpus of base model attempts and retrospections and reaches 48.0%. The online procedure has the highest observed score, while fully offline training outperforms off-policy reflections. This ordering is consistent with a benefit from keeping the attempt and retrospection generators aligned, rather than refreshing atte The same ai evaluation question is explored in Who Holds the Pen? Let Specifications,..., which adds a research perspective. as detailed in the full paper on Arxiv The ai agents story also surfaces in Claude autonomously improved models across 10..., adding another angle.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!