A model can become better at a new task while losing capabilities it already had. Disentangling Self-Distillation examines that trade-off by varying training choices that earlier studies often bundled together. The researchers distinguish acquisition, or accuracy gained on the adaptation task, from retention measured on general benchmarks.
The study investigates privileged-context self-distillation. A teacher sees a reference response in addition to the problem, while a student starts from the same model and must learn to answer without that extra context. The teacher's initial advantage comes from information in its prompt, rather than being a larger or inherently stronger model.
Three choices that can change the result
The researchers vary three axes: whether responses come from the student or the context-conditioned teacher, how closely the teacher follows the updating student, and the direction of the token-level KL loss.
Student rollouts train on prefixes the student itself produces. Teacher rollouts can expose it to informative responses it would not yet generate. Teacher coupling ranges from a frozen copy to an exponential moving average of the student. The loss direction changes how disagreement between their next-token distributions is weighted.
These choices are separable, but they interact with what the new task demands. A rule that directly contradicts a pretrained habit is different from improving a task the model already partly understands.
A controlled adaptation grid
The experimental grid comprises 1,200 self-distillation runs across two models and five tasks. It crosses twenty training configurations with learning rates and seeds, using Qwen2.5-7B-Instruct and Ministral 3 3B Instruct.
Three tasks are ordinary: structured tool calls, chemistry questions, and spatial reasoning. Two are contradictory, including arithmetic in base nine and spatial relations with a rotated compass convention. Each task is adapted separately from the original model.
Acquisition is measured against the model's initial task accuracy. Retention is the change across HellaSwag, MMLU, TruthfulQA-MC2, WinoGrande, and IFEval. Supervised fine-tuning uses the same reference responses as a comparison, although its learning-rate grid is wider. A separate small autoregressive model lets the researchers vary mechanisms under more controlled conditions.
What the authors report
Teacher rollouts help acquisition most when the new task contradicts pretrained behavior, with little retention change in those comparisons. Increasing teacher coupling initially improves acquisition, but the benefit falls past a task- and model-dependent rate. Changing KL direction has a model-dependent retention cost.
The authors also report that supervised fine-tuning reaches the highest acquisition on most tasks while generally paying more in retention. Self-distillation therefore should not be described as winning every adaptation objective.
These findings concern the tested models, tasks, and measurement suite. They do not establish a universal coupling rate or prove preservation of all deployed behavior. For adaptation work, the useful recommendation is to measure task gains and unrelated capability changes separately, then choose training settings for that observed trade-off.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!