Back to AI Research

AI Research

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling... | AI Research

Key Takeaways

  • Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL Long-horizon agents often learn from a simple outcome signal: whether they c...
  • Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards.
  • This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning.
  • To mitigate this problem, we use on-policy distillation (OPD) to provide token-level guidance on the student's own rollouts.
  • We find that the benefit of this guidance depends on the performance gap between the teacher and the student.
Paper AbstractExpand

Reinforcement learning for long-horizon agents typically relies on sparse outcome-based rewards. This leads to a severe cold-start problem, as early-stage policies often fail to solve sampled tasks, leaving little useful reward signal for learning. To mitigate this problem, we use on-policy distillation (OPD) to provide token-level guidance on the student's own rollouts. We find that the benefit of this guidance depends on the performance gap between the teacher and the student. When the teacher substantially outperforms the student, distillation helps guide the student through the early training stage where outcome rewards provide little learning signal. As the gap narrows and eventually reverses, however, continued distillation becomes less beneficial and may hinder further improvement. Motivated by this observation, we propose Gap-Adaptive Teacher Scheduling (GATS), which augments the student's RL objective with an OPD term whose weight adapts to the teacher-student performance gap. Specifically, GATS gradually reduces teacher guidance as the student approaches the teacher's reference performance and withdraws it once that reference is reached. This enables GATS to leverage task-trained teachers smaller than the student, since teacher guidance is primarily needed during early training. Across ALFWorld, WebShop, and ScienceWorld with three Qwen2.5 teacher-student configurations, GATS achieves the highest average success rate among the compared methods in all three configurations, improving over reward-only GRPO by 4.37%-11.87% under matched student rollout budgets. Code is available at this https URL .

Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL

Long-horizon agents often learn from a simple outcome signal: whether they completed a task. But early in training, an agent may fail every attempted task, producing no useful difference between its rollouts. This paper on gap-adaptive teacher scheduling for sparse-reward agentic reinforcement learning addresses that cold-start problem by combining outcome-based reinforcement learning with token-level guidance from a stronger, task-trained teacher. Its central idea is simple: use the teacher heavily when the student is struggling, reduce guidance as the student catches up, and remove the teacher once the student reaches the teacher’s reference performance.

Why sparse rewards create a cold start

The method builds on Group Relative Policy Optimization, or GRPO. For each task, the student generates a group of trajectories, and each trajectory receives an outcome reward based on whether it completes the task. GRPO compares those rewards within the group, assigning positive or negative relative advantages that guide the policy update.
This works poorly when all trajectories fail. If every sampled attempt receives the same reward, the group has no meaningful reward variation. Their relative advantages become zero, so the outcome signal provides little or no policy-gradient information. The student may therefore remain stuck precisely when it needs the most help.
A natural response is to imitate successful teacher trajectories. However, the authors distinguish their approach from ordinary supervised imitation. Instead of training the student on fixed, off-policy expert demonstrations, on-policy distillation (OPD) evaluates the teacher on the student’s own trajectories. At each generated token, the student receives guidance based on the difference between the teacher’s and student’s log-probabilities for that token. This supplies a denser signal while preserving the student’s own interaction context.
The paper’s experiments suggest that OPD is most useful early, when the teacher is substantially better than the student. But its usefulness falls as the gap closes. Near parity, continued teacher guidance can become harmful: it may pull the student toward the teacher’s behavior and restrict improvement beyond the teacher’s capability.

How GATS adapts teacher guidance

Gap-Adaptive Teacher Scheduling, or GATS, combines two objectives: the usual GRPO loss from outcome rewards and the OPD loss from token-level teacher guidance. The OPD term is multiplied by a weight that depends on the measured performance gap.
Before student training begins, the teacher is trained on the target tasks and then frozen. The authors calculate a teacher reference score by averaging its final training measurements over a fixed window. During student training, they similarly track a lagged moving average of the student’s recent training scores. The lag prevents the current update from determining its own teacher weight.
The OPD weight is defined as:
[ \lambda_t = \max\left(1-\frac{M_{S,t}}{M_T}, 0\right), ]
where (M_{S,t}) is the student’s recent score and (M_T) is the teacher reference score. When the student is far below the teacher, the weight is larger. As the student approaches the teacher reference, the weight decreases. Once the student’s moving-average score reaches or exceeds that reference, GATS permanently withdraws the teacher and continues with GRPO alone.
This “guide, then let go” structure is intended to avoid two opposing problems. Withdrawing guidance too early can return the student to the sparse-reward cold start. Keeping it indefinitely can anchor the student near the teacher’s ceiling. GATS instead treats the teacher as early directional support rather than a permanent target.
The approach also permits a teacher smaller than the student. The authors argue that the teacher is mainly needed to provide useful direction during early training, so it does not need to match the student’s eventual scale. They test Qwen2.5 teacher–student pairs of 1.5B → 7B, 1.5B → 14B, and 3B → 7B.

Results across three agent benchmarks

The evaluation covers ALFWorld, WebShop, and ScienceWorld. The reported metrics include in-distribution and out-of-distribution success rates for ALFWorld and ScienceWorld, plus WebShop evaluation performance. Students are trained for 150 updates, with the student rollout budget and common training settings held fixed across methods.
GATS achieves the highest average success rate in all three teacher–student configurations. With a Qwen2.5-1.5B teacher and 7B student, its average success rate is 64.84%, compared with 52.97% for reward-only GRPO—a gain of 11.87 percentage points. With the same teacher and a 14B student, GATS reaches 66.30%, compared with 61.93% for GRPO, a 4.37-point improvement. With a 3B teacher and 7B student, it reaches 64.74%, compared with 52.97%, a 11.77-point improvement.
The authors also report gains on the out-of-distribution splits of ALFWorld and ScienceWorld. In their interpretation, this indicates that the method’s benefit is not limited to reproducing the teacher’s behavior on the training distribution.
The comparisons also illustrate why scheduling matters. Fixed-weight GRPO plus OPD and the SOD baseline, which retains teacher influence throughout training, perform below plain GRPO in some configurations. The paper attributes this to the teacher’s continued influence after it is no longer beneficial. ATOD, which uses a predefined annealing schedule, performs inconsistently: its average success rate ranges from 47.66% to 61.20% across the tested settings. The authors argue that a manually timed schedule can withdraw guidance too early or too late because it does not observe the student’s actual progress.

What the evidence does—and does not—show

The paper’s main empirical claim is that the value of OPD tracks the teacher–student capability gap. In a focused experiment, the authors compare increasingly capable student checkpoints against a fixed teacher. They report that OPD’s gain over GRPO declines as the gap narrows and becomes negative once the student exceeds the teacher. This observation motivates both the adaptive weighting and permanent withdrawal rule.
The reported results support GATS across the three benchmarks and three model configurations tested. They also show that a smaller teacher can assist a larger student under the proposed schedule. However, the evidence is limited to the stated Qwen2.5 configurations, environments, training setup, and rollout budgets. The teacher is trained on the target tasks, and the withdrawal threshold depends on a teacher reference score and moving-average measurements. The paper therefore establishes a particular adaptive strategy in these experiments, rather than proving that the same schedule will work unchanged for every agent, teacher, or environment.
The broader lesson is not that teachers should always guide reinforcement-learning agents. It is that guidance may have a changing role: dense supervision can help an agent discover useful behavior when sparse rewards provide almost no signal, while later progress may require removing that supervision and allowing the student to explore beyond its teacher.

Comments (0)

No comments yet

Be the first to share your thoughts!