Guide, Then Let Go: Gap-Adaptive Teacher Scheduling for Sparse-Reward Agentic RL
Long-horizon agents often learn from a simple outcome signal: whether they completed a task. But early in training, an agent may fail every attempted task, producing no useful difference between its rollouts. This paper on gap-adaptive teacher scheduling for sparse-reward agentic reinforcement learning addresses that cold-start problem by combining outcome-based reinforcement learning with token-level guidance from a stronger, task-trained teacher. Its central idea is simple: use the teacher heavily when the student is struggling, reduce guidance as the student catches up, and remove the teacher once the student reaches the teacher’s reference performance.
Why sparse rewards create a cold start
The method builds on Group Relative Policy Optimization, or GRPO. For each task, the student generates a group of trajectories, and each trajectory receives an outcome reward based on whether it completes the task. GRPO compares those rewards within the group, assigning positive or negative relative advantages that guide the policy update.
This works poorly when all trajectories fail. If every sampled attempt receives the same reward, the group has no meaningful reward variation. Their relative advantages become zero, so the outcome signal provides little or no policy-gradient information. The student may therefore remain stuck precisely when it needs the most help.
A natural response is to imitate successful teacher trajectories. However, the authors distinguish their approach from ordinary supervised imitation. Instead of training the student on fixed, off-policy expert demonstrations, on-policy distillation (OPD) evaluates the teacher on the student’s own trajectories. At each generated token, the student receives guidance based on the difference between the teacher’s and student’s log-probabilities for that token. This supplies a denser signal while preserving the student’s own interaction context.
The paper’s experiments suggest that OPD is most useful early, when the teacher is substantially better than the student. But its usefulness falls as the gap closes. Near parity, continued teacher guidance can become harmful: it may pull the student toward the teacher’s behavior and restrict improvement beyond the teacher’s capability.
How GATS adapts teacher guidance
Gap-Adaptive Teacher Scheduling, or GATS, combines two objectives: the usual GRPO loss from outcome rewards and the OPD loss from token-level teacher guidance. The OPD term is multiplied by a weight that depends on the measured performance gap.
Before student training begins, the teacher is trained on the target tasks and then frozen. The authors calculate a teacher reference score by averaging its final training measurements over a fixed window. During student training, they similarly track a lagged moving average of the student’s recent training scores. The lag prevents the current update from determining its own teacher weight.
The OPD weight is defined as:
[ \lambda_t = \max\left(1-\frac{M_{S,t}}{M_T}, 0\right), ]
where (M_{S,t}) is the student’s recent score and (M_T) is the teacher reference score. When the student is far below the teacher, the weight is larger. As the student approaches the teacher reference, the weight decreases. Once the student’s moving-average score reaches or exceeds that reference, GATS permanently withdraws the teacher and continues with GRPO alone.
This “guide, then let go” structure is intended to avoid two opposing problems. Withdrawing guidance too early can return the student to the sparse-reward cold start. Keeping it indefinitely can anchor the student near the teacher’s ceiling. GATS instead treats the teacher as early directional support rather than a permanent target.
The approach also permits a teacher smaller than the student. The authors argue that the teacher is mainly needed to provide useful direction during early training, so it does not need to match the student’s eventual scale. They test Qwen2.5 teacher–student pairs of 1.5B → 7B, 1.5B → 14B, and 3B → 7B.
Results across three agent benchmarks
The evaluation covers ALFWorld, WebShop, and ScienceWorld. The reported metrics include in-distribution and out-of-distribution success rates for ALFWorld and ScienceWorld, plus WebShop evaluation performance. Students are trained for 150 updates, with the student rollout budget and common training settings held fixed across methods.
GATS achieves the highest average success rate in all three teacher–student configurations. With a Qwen2.5-1.5B teacher and 7B student, its average success rate is 64.84%, compared with 52.97% for reward-only GRPO—a gain of 11.87 percentage points. With the same teacher and a 14B student, GATS reaches 66.30%, compared with 61.93% for GRPO, a 4.37-point improvement. With a 3B teacher and 7B student, it reaches 64.74%, compared with 52.97%, a 11.77-point improvement.
The authors also report gains on the out-of-distribution splits of ALFWorld and ScienceWorld. In their interpretation, this indicates that the method’s benefit is not limited to reproducing the teacher’s behavior on the training distribution.
The comparisons also illustrate why scheduling matters. Fixed-weight GRPO plus OPD and the SOD baseline, which retains teacher influence throughout training, perform below plain GRPO in some configurations. The paper attributes this to the teacher’s continued influence after it is no longer beneficial. ATOD, which uses a predefined annealing schedule, performs inconsistently: its average success rate ranges from 47.66% to 61.20% across the tested settings. The authors argue that a manually timed schedule can withdraw guidance too early or too late because it does not observe the student’s actual progress.
What the evidence does—and does not—show
The paper’s main empirical claim is that the value of OPD tracks the teacher–student capability gap. In a focused experiment, the authors compare increasingly capable student checkpoints against a fixed teacher. They report that OPD’s gain over GRPO declines as the gap narrows and becomes negative once the student exceeds the teacher. This observation motivates both the adaptive weighting and permanent withdrawal rule.
The reported results support GATS across the three benchmarks and three model configurations tested. They also show that a smaller teacher can assist a larger student under the proposed schedule. However, the evidence is limited to the stated Qwen2.5 configurations, environments, training setup, and rollout budgets. The teacher is trained on the target tasks, and the withdrawal threshold depends on a teacher reference score and moving-average measurements. The paper therefore establishes a particular adaptive strategy in these experiments, rather than proving that the same schedule will work unchanged for every agent, teacher, or environment.
The broader lesson is not that teachers should always guide reinforcement-learning agents. It is that guidance may have a changing role: dense supervision can help an agent discover useful behavior when sparse rewards provide almost no signal, while later progress may require removing that supervision and allowing the student to explore beyond its teacher.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!