Back to AI Research

AI Research

Instruction-tuning transfer can help in one direction and hurt in the other

Key Takeaways

  • A signed transfer map predicts which task types support a held-out target under a fixed training budget.
  • The study finds that more sources and greater similarity are unreliable sel
  • The study finds that more sources and greater similarity are unreliable selection rules.
  • Adding another instruction-tuning task does not give a model free extra capacity.
  • Under a fixed example budget, its examples displace examples from other tasks.

Adding another instruction-tuning task does not give a model free extra capacity. Under a fixed example budget, its examples displace examples from other tasks. A Helps B While B Hurts A measures those trade-offs and finds that transfer between task types can be both negative and directional.

Measuring benefit instead of resemblance

The researchers generate eight task types from one neuroscience corpus, including definitions, causal explanations, multi-hop questions and methodological critiques. A transfer map estimates how each source task changes accuracy on a target task excluded from training.
Keeping the corpus fixed helps separate the task being asked from differences between unrelated datasets. The main corpus contains 110 papers, split by publication date into 88 training papers and 22 evaluation papers. A cancer-immunotherapy abstract corpus provides a separate corpus-shift control.
The study spans 751 fine-tuning runs across Qwen3 sizes and Mistral-24B. Each run uses 6,000 examples, divided equally among its included sources. The map therefore estimates effects inside fixed-budget mixtures; it is not a universal score for whether a task is valuable.

Why a symmetric similarity score is insufficient

On Qwen3-32B, replacing an average source with definitions improves multi-hop accuracy by about 5.3 percentage points. Adding multi-hop examples to a mixture instead lowers definition accuracy by about 1.9 points in absolute terms. These are different directions and different effect definitions, not two measurements of one symmetric relationship.
Methodological critiques are a particularly harmful source for multi-hop questions in the measured maps. That does not mean critique training is generally undesirable. A source can interfere with one target while serving another, and excluding it may sacrifice capabilities outside the chosen objective.
The authors also show why a positive association between mixture size and accuracy cannot justify including every task. Mixture size is the sum of the source-inclusion indicators. A larger mixture may perform better because its average additional source helps, while still containing a source that hurts.

Predicting unseen mixtures before training

The map selects sources and records predicted gains before the corresponding test runs. Across eight non-vacuous test cells, its mean absolute prediction error is 2.31 percentage points, compared with 6.20 for a baseline predicting no mixture effect.
Six measured effects are gains, reaching 13.76 points over three seeds. In two cells, small predicted improvements become small negative effects whose confidence intervals include zero. The method therefore does not identify a winning subset without uncertainty.
The evaluation's open-answer tasks use an embedding-similarity scoring rule. That operational definition of accuracy should remain visible when interpreting reasoning gains; similarity to a reference is not an independent factual audit.
The paper reports transfer across tested model sizes and families, but also a changed map on the second corpus. Several properties change together in that control, including document granularity and training-pool size. For a new domain, these results favor measuring source–target effects under the actual budget rather than importing the published map as a ready-made recipe.

Comments (0)

No comments yet

Be the first to share your thoughts!