Back to AI Research

AI Research

AdviSD: Learning to Advise Frontier LLMs via Target... | AI Research

Key Takeaways

  • What the paper is about A small trainable advisor can steer a frozen language-model executor using natural-language advice.
  • A small trainable advisor can steer a frozen language-model executor using natural-language advice.
  • In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice.
  • However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts.
  • In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections.
Paper AbstractExpand

A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.

What the paper is about

A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule. The same ai evaluation question is explored in Verifiable Visual Rewards Transfer from Synthetic..., which adds a research perspective.

What it covers

\uselogo AdviSD : Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation Rishabh Agrawal Affiliation: \thepa Hejie Cui Affiliation: \thepa Shasha Li Affiliation: \thepa Shanchan Wu Affiliation: \thepa Sercan Ö. Arık Affiliation: \thepa Affiliation: University of Southern California Abstract A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor’s future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model’s eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation ( AdviSD ), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2–6.4 percentage points on BFCL-v3 and by 3.9–5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule. 1 Introduction Frontier language models are usually served through APIs that accept queries but do not let users change the model weights. When such a model acts as an agent, a smaller trainable advisor can adapt it from the outside: the advisor observes the interaction and recommends what the frozen agent, which we call the executor, should do next ( Li et al., 2023 ; Li et al., 2025 ) . The advisor can be trained with reinforcement learning on the rewards of completed interactions ( Asawa et al., 2026 ) . However, an episode reward summarizes task performance without specifying which of the advisor’s decisions should have been different, or how. Completed interactions contain more specific evidence: executor responses, tool results, and task checks. These observations and the episode reward form the feedback that reflection uses to propose revisions ( Liu et al., 2026 ; Yeo et al., 2026 ) . In feedback-conditioned self-distillation, a teacher that sees this feedback supervises a student that sees only the original context ( Hübotter et al., 2026 ; Agrawal et al., 2026b ) . For an advisor, however, the learned advice acts through another model: revising it need not change what the executor does. Consider an executor that searches reliably for suitable flights without guidance but sometimes guesses the passenger identifier when booking. After a reservation fails because that identifier is wrong, reflection may propose looking up the passenger and using the returned identifier in the booking. It may also propose more detailed advice for the earlier flight search. Both revisions are valid, but their usefulness differs for this executor: the booking revision addresses the error, whereas the search revision elaborates a procedure the executor already follows unaided. The teacher’s supervision can reflect both revisions, and learning from the redundant search revision can still change the advisor’s shared parameters and affect advice elsewhere, for better or worse. Hence, we ask which revisions provide useful supervision for a given advisor–executor pair. Our contributions are as follows: A theoretical account of correction selection. Lemma 1 separates agreeing with a feedback-conditioned teacher from improving execution. We then study repeated learning in a shared-parameter model where the teacher’s preferences and the executor’s behavior under any given advice stay fixed during training. As advice improves, preventable failures become less frequent, while failures unaffected by advice persist. If corrections from these persistent failures teach a weaker preference for useful advice, their growing share of supervision limits learning. Retaining them less often than other corrections raises the performance the advisor eventually reaches, whether or not reward learning is added (Theorems 1 – 2 ). By contrast, randomly discarding corrections at the same rate across both types leaves eventual performance unchanged when learning only from corrections, but can improve it alongside reward learning (Corollary 1 ). Therefore, our matched-count random control tests whether targeted selection helps beyond reducing supervision (Section 2 ). Figure 1: Overview of AdviSD . (A) Multi-turn advising. Before each executor response, the advisor provides advice or abstains; the frozen executor uses its native context and any issued advice. (B) Training from targeted feedback. Reflection proposes corrections at advice decisions in imperfect episodes. Among flagged decisions, AdviSD retains original abstentions and those whose scores with and without issued advice differ by more than a calibrated threshold. The pre-update advisor computes both on the same recorded executor response. At retained decisions, a feedback-conditioned copy of the pre-update advisor teaches the trainable advisor, which sees only the original context. This self-distillation supplements GRPO on all episodes. AdviSD : multi-turn advising with targeted feedback. AdviSD separates proposing revisions from choosing where to learn (Figure 1 ). The advisor scores the same recorded executor response under two contexts, one containing its issued advice and one without it, so selection requires neither executor likelihoods nor additional executor rollouts. We use the magnitude of the score difference as a predictive signal for selecting decisions to supervise. At selected decisions, a feedback-conditioned copy of the pre-update advisor supervises the trainable advisor. This targeted self-distillation complements outcome-based GRPO ( Shao et al., 2024 ) . Reflection and scoring are used only during training. At deployment, the advisor decides before each executor turn whether to advise or abstain. Empirical results. With Qwen3-8B advisors for Gemini and Claude, AdviSD has the highest in-domain aggregates on BFCL-v3 and EnvScaler among the compared methods. It exceeds advisor-GRPO by 4.2–6.4 percentage points on BFCL-v3 and 3.9–5.1 score points on EnvScaler. Across these settings, its 2.5–4.9-point advantage over matched-count random selection supports choosing which decisions to supervise beyond reducing supervision. Without retraining, AdviSD improves out-of-domain macro-averages by 2.7–3.6 points over standalone execution. It also transfers across executor versions and model families and exceeds transferred GRPO by 3.1 percentage points in both cross-family directions (Section 7 ). 2 Related Work Advising and prompt optimization. Frozen models can be adapted through reusable instructions or context-dependent guidance. GEPA uses reflection to optimize reusable instructions ( Agrawal et al., 2026a ) , while Directional Stimulus Prompting, Matryoshka Pilot, and Advisor Models train smaller models to guide frozen ones ( Li et al., 2023 ; Li et al., 2025 ; Asawa et al., 2026 ) . Self-Refine and Reflexion use feedback to revise outputs or guide later attempts without updating model weights ( Madaan et al., 2023 ; Shinn et al., 2023 ) . Our advisor-GRPO baseline follows Advisor Models’ outcome-based training through a response-level tool-use interface; AdviSD adds targeted feedback-conditioned self-distillation. Feedback-conditioned distillation. Feedback-conditioned on-policy self-distillation uses additional information to supervise a policy on its own trajectories. SDPO obtains this information from environment feedback or successful rollouts ( Hübotter et al., 2026 ) , and DistIL optimizes forward cross-entropy with sequence-level credit assignment ( Agrawal et al., 2026b ) . Other approaches focus on constructing and allocating supervision. HERO constructs turn-level feedback, and HinT-SD selects failure-relevant action spans for self-distillation ( Liu et al., 2026 ; Yeo et al., 2026 ) . LOPD builds teacher context from retrieved experience, and DART-SD retrieves references for recovery ( Zhang et al., 2026 ; Xu et al., 2026 ) . SAGE-OPD selects and weights teacher supervision at individual turns ( Zhou et al., 2026 ) . In contrast, AdviSD addresses which corrections to learn when the trained policy advises a separate executor rather than directly performing the task. Predictive contrasts and selection. Comparing predictions made with different information can yield a learning signal. RLCSD contrasts correct and incorrect hints, and OCSD compares full and observation-ablated contexts ( Pan et al., 2026 ; Yang et al., 2026b ) . PBSD and RLSD use paired predictions to refine turn-level credit or token updates ( Tian et al., 2026 ; Yang et al., 2026a ) . AdviSD applies paired scoring to a separate executor’s recorded response, with and without the issued advice. It uses the contrast magnitude to select auxiliary supervision for an advisor whose advice acts through a frozen executor, while leaving the rollout batch’s GRPO advantages unchanged. Appendix H provides the full discussion and further comparisons. 3 Background and Problem Formulation Advising a frozen executor. Before each executor response k k , the advisor reads a context g k g_{k} (the visible interaction, tool schemas, and its earlier advice) and samples an action a k ∼ π θ ( ⋅ ∣ g k ) a_{k}\sim\pi_{\theta}(\cdot\mid g_{k}) , which is either advice text or the abstention sequence . The frozen executor responds with m k ∼ ρ ( ⋅ ∣ h k , e ( a k ) ) m_{k}\sim\rho(\cdot\mid h_{k},e(a_{k})) , where h k h_{k} is its native history and e ⁡ ( a k ) e(a_{k}) is the advice text, or an empty string if the advisor abstains. Advice goes into a temporary request rather than the executor’s persistent history. The advisor makes a new decision before every response, including text-only responses and those that follow tool results; parallel tool calls within one response share a single decision. An episode ends with reward R ∈ [ 0 , 1 ] R\in[0,1] . Two learning signals. For n n rollouts of a task, GRPO assigns episode i i the advantage A ^ i = ( R i − R ¯ ) / ( s R + ϵ A ) \widehat{A}{i}=(R{i}-\bar{R})/(s_{R}+\epsilon_{A}) , where R ¯ \bar{R} and s R s_{R} are the group’s reward mean and standard deviation, and ϵ A > 0 \epsilon_{A}>0 stabilizes the denominator. This advantage applies to every advisor token generated in the episode, including abstentions. We denote the clipped GRPO loss with reference-policy regularization by ℒ base \mathcal{L}{\rm base} ( Shao et al., 2024 ) . Feedback-conditioned self-distillation supervises individual advice decisions. The teacher , a copy of the pre-update advisor with parameters θ ¯ \bar{\theta} , sees the original context augmented with feedback from the completed interaction. The trainable student sees only the original context and learns to match the teacher’s next-token distributions, which are held fixed during optimization. AdviSD chooses which decisions to supervise, including originally abstaining decisions flagged by reflection, and leaves the GRPO advantages unchanged. 4 Why the Choice of Corrections Matters We examine why fitting a teacher need not improve execution (Section 4.1 ), then show how the retained corrections determine the learning limit in a shared-parameter model (Section 4.2 ). 4.1 A single update through the executor Fix an interaction state and advice prefix h h , and let p θ p{\theta} and q h q_{h} be positive student and teacher distributions on a fixed finite token set S h S_{h} , possibly the full vocabulary. The student is differentiable near the pre-update parameters θ ¯ \bar{\theta} . Choosing token v v , completing the advice with the pre-update advisor, and running the executor induces an execution law K h , v K_{h,v} over responses and outcomes, excluding advice text. For a bounded task score W W , define V h ​ ( v ) = 𝔼 Z ∼ K h , v ​ [ W ⁡ ( Z ) ] , J h ​ ( θ ) = ∑ v ∈ S h p θ ​ ( v ) ​ V h ​ ( v ) . V_{h}(v)=\mathbb{E}{Z\sim K{h,v}}[W(Z)],\qquad J_{h}(\theta)=\sum_{v\in S_{h}}p_{\theta}(v)V_{h}(v). These are the expected score after choosing v v and its average under the student. Only next-token probabilities vary during differentiation; the teacher, support, completion policy, and execution laws remain fixed. Feedback does not directly provide each token’s expected execution value. We therefore compare q h q_{h} with a normalized reference q h α ​ ( v ) ∝ p θ ¯ ​ ( v ) ​ e α ​ V h ​ ( v ) q_{h}^{\alpha}(v)\propto p_{\bar{\theta}}(v)e^{\alpha V_{h}(v)} , α > 0 \alpha>0 , which reweights the student toward higher-value tokens. This value-tilted teacher ( Peters et al., 2010 ) is an analytical benchmark; AdviSD does not construct it. Lemma 1 (Teacher fitting versus execution improvement) . Under this setup, let ϕ h ​ ( v ) = ∇ θ ​ log ​ p θ ​ ( v ) | θ ¯ \phi_{h}(v)=\nabla_{\theta}\log p_{\theta}(v)|{\bar{\theta}} and g h = ∇ θ KL ( p θ ∥ q h ) | θ ¯ g{h}=\nabla_{\theta}\operatorname{KL}(p_{\theta}|q_{h})|{\bar{\theta}} . For every fixed α > 0 \alpha>0 , g h = − α ∇ J h ( θ ¯ ) + e h , e h = 𝔼 v ∼ p θ ¯ [ ϕ h ( v ) log q h α ​ ( v ) q h ​ ( v ) ] . g{h}=-\alpha\nabla J_{h}(\bar{\theta})+e_{h},\qquad e_{h}=\mathbb{E}{v\sim p{\bar{\theta}}}\left[\phi_{h}(v)\log\frac{q_{h}^{\alpha}(v)}{q_{h}(v)}\right]. (1) Reverse-KL descent at θ ¯ \bar{\theta} toward the reference follows α ∇ J h \alpha\nabla J_{h} , whereas descent toward the actual teacher follows α ∇ J h − e h \alpha\nabla J_{h}-e_{h} , whose residual can reinforce or oppose value ascent. At an insensitive prefix, every supported token induces the same execution law, so ∇ J h = 0 \nabla J_{h}=0 and the reference equals the student. Changing token probabilities cannot improve this local objective, but teacher fitting can still change advice elsewhere through shared parameters, for better or worse. Teacher agreement alone does not tell us which. Appendix B gives the full proof and extensions. 4.2 Repeated updates and the learning limit Lemma 1 concerns one update. We now introduce a simplified shared-parameter model to study how repeated learning changes which failures occur and which corrections supply supervision. Setup. The advisor chooses between advice 1 and advice 2, selecting advice 1 with probability p ⁡ ( θ ) = 1 / ( 1 + e − θ ) p(\theta)=1/(1+e^{-\theta}) . A single log-odds parameter θ \theta is shared across a fixed mixture of sensitive ( S S ) and insensitive ( I I ) situations, each with positive probability. In sensitive situations, advice 1 succeeds more often than advice 2; in insensitive situations, both induce the same execution law. These laws remain fixed, with success probabilities in ( 0 , 1 ) (0,1) , so expected success J ⁡ ( θ ) J(\theta) increases with θ \theta . A failure of type j ∈ { I , S } j\in{I,S} supplies a fixed teacher Q j Q_{j} , positive on both advice choices, with target log-odds ℓ j = log ⁡ [ Q j ​ ( 1 ) / Q j ​ ( 2 ) ] \ell_{j}=\log[Q_{j}(1)/Q_{j}(2)] . Even when both advice choices lead to identical executor behavior, the teacher may prefer one over the other. We separately assume ℓ I 1 / 2 p>1/2 , learning from Q I Q_{I} pushes that probability down toward 1 / 2 1/2 . Because θ \theta is shared, this also makes advice 1 less likely in sensitive situations, where it succeeds more often. Thus, a neutral teacher can weaken useful advice elsewhere without favoring advice 2. Each episode contains a fixed positive number of independent, identically distributed (situation, advice, outcome) samples from this model. Failures supply correction proposals up to a fixed positive cap, with uniform subsampling if the cap is exceeded. Proposals of type j j are retained independently with fixed probability r j ∈ ( 0 , 1 ] r_{j}\in(0,1] ; no gating means r I = r S = 1 r_{I}=r_{S}=1 . The episode loss averages student-to-teacher reverse KL over retained corrections and is zero if none remain. Samples and selections are held fixed during differentiation. Retained supervision. Let B ⁡ ( θ ) B(\theta) be the probability that an episode retains any correction, and let ω I ​ ( θ ) \omega_{I}(\theta) be the expected fraction of insensitive corrections conditional on retaining at least one. The mean target log-odds is μ ⁡ ( θ ) = ω I ​ ( θ ) ​ ℓ I + [ 1 − ω I ​ ( θ ) ] ​ ℓ S \mu(\theta)=\omega_{I}(\theta)\ell_{I}+[1-\omega_{I}(\theta)]\ell_{S} . The quantity B B measures exposure , how often episodes receive supervision, while μ \mu summarizes the composition of that supervision, the mixture of teacher targets. As θ \theta increases, sensitive failures become less frequent while the insensitive failure rate stays fixed. The insensitive teacher therefore receives a growing share of supervision, so μ \mu decreases. Theorem 1 (The retained mixture sets the learning limit) . Under this setup, with distillation alone, the expected gradient of the sampled episode loss is g ⁡ ( θ ) = B ⁡ ( θ ) ​ p ​ ( 1 − p ) ​ [ θ − μ ⁡ ( θ ) ] . g(\theta)=B(\theta),p(1-p),[\theta-\mu(\theta)]. (2) The continuous-time update θ ˙ = − g ⁡ ( θ ) \dot{\theta}=-g(\theta) converges from every finite initialization to a unique equilibrium θ ∗ = μ ⁡ ( θ ∗ ) ∈ ( ℓ I , ℓ S ) \theta^{}=\mu(\theta^{})\in(\ell_{I},\ell_{S}) . Decreasing r I / r S r_{I}/r_{S} strictly increases both θ ∗ \theta^{} and J ⁡ ( θ ∗ ) J(\theta^{}) . Changing the episode size or proposal cap, or scaling both retention probabilities by the same admissible positive factor, leaves the limit unchanged. Each teacher contributes p ⁡ ( 1 − p ) ​ ( θ − ℓ j ) p(1-p)(\theta-\ell_{j}) to the gradient; averaging yields Eq. ( 2 ). Learning settles where the advisor’s log-odds equal the retained teachers’ mean target. Retaining insensitive corrections less often than sensitive ones raises this target and the resulting learning limit. Independently thinning both types at the same rate changes exposure but not the target. Appendices C.1 – C.3 provide the proof, scope, and extensions. This model retains corrections by situation type. AdviSD instead uses an observable predictive contrast (Section 5 ), whose usefulness we evaluate through ablations (Section 2 ). When reward learning is added, exposure can also affect eventual performance (Section 6 ), so the ablations include a matched-count random control. 5 AdviSD : Advisor Self-Distillation AdviSD trains only the advisor, combining outcome-based GRPO with targeted self-distillation (Figure 1 ). An external reflector proposes corrections, a predictive selector chooses decisions to supervise, and a feedback-conditioned pre-update advisor teaches a student that sees only the original context. These stages run only during training (Algorithm 1 ; Appendix D ). 5.1 Reflection proposes corrections For each eligible imperfect episode i i , the reflector uses executor responses, tool outcomes, and checks to flag at most b refl b_{\rm refl} advice decisions 𝒥 i \mathcal{J}{i} with correction feedback. Later events can explain failures, but proposed advice uses information available at the original decision (Appendices F and D ). 5.2 A paired score selects where to learn Scoring. Let y k = ( y k , 1 , … , y k , T k ) y{k}=(y_{k,1},\ldots,y_{k,T_{k}}) denote the recorded executor response, serialized and tokenized with the advisor’s tokenizer. It includes tool calls in their recorded order but excludes subsequent tool results. We construct two scoring contexts, C k + C_{k}^{+} and C k − C_{k}^{-} , from the request sent to the executor rather than from the advisor’s context g k g_{k} . Both contain the same pre-response history and tool schemas and differ only in the issued advice, which C k + C_{k}^{+} includes and C k − C_{k}^{-} omits. The pre-update advisor scores each token of y k y_{k} under both contexts: c k = 1 T k ​ ∑ t = 1 T k [ log ⁡ π θ ¯ ​ ( y k , t ∣ C k + , y k , ϵ c |c_{i,k}|>\epsilon_{c} . An original abstention has C k + = C k − C_{k}^{+}=C_{k}^{-} and hence c k = 0 c_{k}=0 , so the contrast cannot detect missed advice; flagged abstentions therefore bypass scoring. A proposal to abstain after issued advice must still pass the numeric gate. 5.3 Self-distillation from targeted feedback At each retained decision, ζ i , k \zeta_{i,k} combines local execution evidence, relevant checks, episode score, and reflection feedback. The teacher sees it prepended to the context g i , k g_{i,k} ; the student sees only g i , k g_{i,k} . Both predict along the originally sampled advice a i , k a_{i,k} , including : q i , k , t = sg π θ ¯ ( ⋅ ∣ ζ i , k ⊕ g i , k , a i , k , 0 . \dot{\theta}{j}=F{j}(\theta_{j}),\qquad F_{j}(\theta)=c_{0}J^{\prime}(\theta)-\lambda g_{j}(\theta),\qquad c_{0}\geq 0,\quad\lambda>0. Here J ′ ​ ( θ ) J^{\prime}(\theta) is the gradient of expected success and g j ​ ( θ ) g_{j}(\theta) is the expected distillation gradient under retention rule j j ; c 0 c_{0} and λ \lambda weight the two learning signals. The reward term idealizes finite-step GRPO–AdamW training as exact gradient ascent. Let a 0 a_{0} denote the distillation-only equilibrium without gating (Theorem 1 ). Theorem 2 (A higher learning limit under a common reward objective) . Under the preceding two-teacher model, suppose G G independently retains insensitive and sensitive proposals with fixed probabilities r I G , r S G ∈ ( 0 , 1 ] r_{I}^{G},r_{S}^{G}\in(0,1] , respectively, where r I G θ 0 ∞ , J ⁡ ( θ G ∞ ) > J ⁡ ( θ 0 ∞ ) . \theta_{G}^{\infty}>\theta_{0}^{\infty},\qquad J(\theta_{G}^{\infty})>J(\theta_{0}^{\infty}). (6) These inequalities hold even when the dynamics have multiple equilibria. If the common initialization is at or above a 0 a_{0} , then θ G ​ ( t ) > θ 0 ​ ( t ) \theta_{G}(t)>\theta_{0}(t) for every t > 0 t>0 . Selection changes both which corrections are learned and how often learning from corrections occurs . The mean retained teacher target μ j \mu_{j} captures the first, and the probability B j B_{j} that an episode receives supervision captures the second (Section 4.2 ). At the same parameter value θ \theta , the difference between the two updates is F G − F 0 λ ​ p ​ ( 1 − p ) = B G ​ ( μ G − μ 0 ) ⏟ composition + ( B 0 − B G ) ​ ( θ − μ 0 ) ⏟ exposure . \frac{F_{G}-F_{0}}{\lambda p(1-p)}=\underbrace{B_{G}(\mu_{G}-\mu_{0})}{\text{composition}}+\underbrace{(B{0}-B_{G})(\theta-\mu_{0})}{\text{exposure}}. (7) The composition term is positive because selection gives more weight to the sensitive teacher, which assigns a higher probability to advice 1 than the insensitive teacher. The exposure term comes from less frequent supervision, and its sign depends on θ \theta . Above a 0 a{0} , ungated distillation pulls θ \theta downward, so reducing this pull helps. Below a 0 a_{0} , distillation pushes θ \theta upward, so reducing supervision can slow early progress. A higher learning limit therefore need not mean faster learning from the start. Both dynamics move upward below a 0 a_{0} , and F G > F 0 F_{G}>F_{0} at and above it. Each trajectory remains bounded and cannot cross an equilibrium of its own dynamics; together, these properties establish the ordered limits. Appendices C.2 and C.3 provide the proof, an example of slower initial learning, and a stationary random-teacher extension. For c 0 > 0 c_{0}>0 , reward-only ascent approaches the model’s best achievable success. A fixed λ > 0 \lambda>0 instead produces a finite balance between reward ascent and teacher fitting, and at that balance selection gives higher success than no gating. The theorem compares the two combined rules with each other; Section 7 tests gains over outcome-only GRPO under finite training budgets. Because selection also changes how often supervision occurs, outperforming no gating does not by itself show that choosing particular corrections helps. Retaining every proposal independently with the same fixed probability preserves the teacher mixture and the distillation-only limit, yet can improve eventual performance when reward learning is present (Corollary 1 ). This motivates our matched-count random control (Section 2 ). On each of the control’s own episodes, a shadow AdviSD gate determines how many feasible, reflection-flagged issued-advice decisions to retain. The control then randomly selects the same number of eligible decisions, giving each decision an equal chance of being chosen. Unlike independent thinning, these episode-dependent quotas need not preserve the ungated teacher mixture (Appendix E ). What donor calibration controls. Donor calibration s The google story also surfaces in Google’s Gemini AI Accessed Three Outside..., adding another angle. as detailed in the full paper on Arxiv The google story also surfaces in Google AI Releases TimesFM 3 for..., adding another angle.

Comments (0)

No comments yet

Be the first to share your thoughts!