An agent can make a wrong move early in a task and spend the remaining turns failing to recover. PivotOPD studies that problem in multi-turn language agents, where an action changes the environment and therefore changes the decisions the model faces next. The proposed training method teaches both how to avoid a decisive mistake and how to continue from the state that mistake creates.
Find the turn that changes the outcome
The authors define a pivotal mistake as an action that lengthens the shortest remaining route to completing a task, or makes the task unsolvable. In ALFWorld, a household-task environment with a symbolic oracle, they can measure that remaining route after each action. Their preliminary analysis covers Qwen3 models ranging from 8B to 235B parameters.
Across those models, 59% of failed trajectories contain a pivotal mistake. The first such turn typically arrives early, with a median between turns eight and twelve in a 30-turn attempt. In replays of failed Qwen3-8B trajectories, correcting that first pivotal action raises success from 8% to 59%. Keeping the mistake but forcing the oracle's actions for the next two turns reaches 58%. These are counterfactual replay results with oracle guidance, not the autonomous success rate of a deployed agent.
The distinction matters for training. A model that rarely discovers a recovery on its own receives few examples of what successful recovery looks like, even when the environment still permits it.
Teach prevention and recovery separately
Most environments lack an oracle that identifies the shortest route from every state. PivotOPD instead asks a teacher model to inspect a trajectory and propose candidate mistake turns with corrected actions. A candidate becomes pivotal for training when the student's committed action disagrees with the teacher's proposed action.
A frozen copy of the student receives those actions as hints and supplies token-level training targets. Preventive distillation uses a reverse-KL objective to move the student away from the original mistake. Recovery distillation uses forward KL on responses generated with recovery hints, then trains the student without exposing those hints. The researchers combine both signals with group-based reinforcement learning.
The recovery examples begin after the mistaken action. Later recovery turns use a copied environment that replays preceding actions, so the method teaches a continuation from the resulting state rather than pretending the error never happened.
Read the benchmark gains within their scope
The study compares thirteen baselines on ALFWorld, WebShop and search-based question answering, using Qwen3-1.7B and Qwen3-8B students. The authors report the strongest average performance for PivotOPD with both students. Results in the main table average three seeds, with teacher models held consistent across methods for each student. They also report improved issue resolution with a Nemotron-3.5 student on SWE-Bench Verified.
Teacher judgment remains part of the method: on failed ALFWorld trajectories, a detected pivotal turn falls within one turn of the oracle-labeled turn in 77.8% of cases on average. That is useful agreement, but it leaves room for incorrect diagnoses. The paper supports targeted recovery training on the tested tasks; it does not establish that every consequential error is recoverable or that teacher-selected corrections are always right.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!