What the paper is about
Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: \textit{confidence}. Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or early-stopping mechanism. Despite this, self-supervised confidence fine-tuning makes reasoning more efficient, reducing generated tokens by up to 25% at matched accuracy across Gemma, Qwen, Nemotron, and GPT-OSS models on mathematical, scientific, and coding reasoning benchmarks, with efficiency gains comparable to methods that explicitly optimize for shorter reasoning. Analysis of reasoning episodes further shows that confidence supervision largely preserves the base models' high-level reasoning composition rather than selectively suppressing particular behaviors. Our results suggest that efficient reasoning may emerge as a downstream consequence of learning metacognitive signals, without being directly optimized. The nvidia story also surfaces in NVIDIA Launches Cosmos 3 Edge for..., adding another angle.
What it covers
Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency Parsa Hosseini † † thanks: Correspondence to: [email protected] Affiliation: University of Maryland Affiliation: AI Foundations, Capital One Akasha Tigalappanavara Affiliation: AI Foundations, Capital One Sumit Nawathe Affiliation: University of Maryland Chenrui Fan Affiliation: University of Maryland Sourya Basu Genta Indra Winata Anirban Das Soheil Feizi Nima Chitsazan Affiliation: University of Maryland Affiliation: AI Foundations, Capital One Abstract Reasoning models often generate very long reasoning traces, making inference computationally expensive. Existing approaches typically improve efficiency either through inference-time early-stopping mechanisms or by explicitly encouraging shorter reasoning during training, for example through reinforcement learning with length penalties. We show that substantial efficiency gains can instead emerge from a different kind of supervision: confidence . Using a self-supervised procedure, we fine-tune reasoning models to predict their confidence in the answer at intermediate points along their own reasoning trajectories using only 600 training problems. Confidence is used only as a training target: the loss contains no objective for reasoning length, efficiency, or stopping. At inference, the fine-tuned models use the standard generation procedure, with no confidence elicitation or early-stopping mechanism. Despite this, self-supervised confidence fine-tuning makes reasoning more efficient, reducing generated tokens by up to 25% at matched accuracy across Gemma, Qwen, Nemotron, and GPT-OSS models on mathematical, scientific, and coding reasoning benchmarks, with efficiency gains comparable to methods that explicitly optimize for shorter reasoning. Analysis of reasoning episodes further shows that confidence supervision largely preserves the base models’ high-level reasoning composition rather than selectively suppressing particular behaviors. Our results suggest that efficient reasoning may emerge as a downstream consequence of learning metacognitive signals, without being directly optimized. 1 Introduction Reasoning language models have achieved strong performance across a wide range of challenging tasks, but often at substantial inference cost. Modern reasoning models can generate tens of thousands of tokens before producing an answer, motivating growing interest in making reasoning more efficient. Efficient Reasoning Existing approaches Teach efficiency directly Our approach Teach confidence, not efficiency Length penalty RL CoT compression Early truncation Reasoning model Reasoning model Confidence Supervision self-supervised training no efficiency objective Figure 1: Teaching confidence instead of efficiency. Existing methods explicitly optimize for shorter reasoning traces; we supervise only confidence. Several lines of work have sought to improve the efficiency of reasoning in LLMs. Despite their differences, these approaches all make efficiency an explicit part of either training or inference. At inference time, early-stopping methods monitor signals such as confidence ( Yang et al., 2025 ; Hosseini et al., 2026 ) , uncertainty ( Wang et al., 2025 ) , or answer stability ( Liu and Wang, 2025 ) and terminate reasoning when further computation appears unnecessary. At training time, other approaches directly encourage shorter reasoning, for example through reinforcement learning with length penalties ( Liu et al., 2025 ; Arora and Zanette, 2025 ) or fine-tuning on selected concise reasoning trajectories ( Zhao et al., 2026 ; Munkhbat et al., 2025 ) . These approaches either directly control when reasoning stops or explicitly teach the model to produce shorter reasoning. In this work, we study a different question: Can efficient reasoning emerge without directly training the model to produce shorter reasoning traces? We introduce ConfSFT , a self-supervised fine-tuning procedure that trains a reasoning model to predict its confidence at intermediate points along its own reasoning trajectories. The confidence targets are derived entirely from the model’s own token probabilities and require neither gold answers nor external judges. Crucially, confidence is the only supervised signal: reasoning tokens are masked from the loss, and the objective contains no term for reasoning length, efficiency, or stopping. At inference, the fine-tuned model uses standard generation, with no confidence elicitation or stopping mechanism. Figure 1 illustrates this. The motivation for this supervision comes from a simple observation: confidence provides a meaningful signal about the model’s intermediate reasoning state. Across intermediate states, higher confidence is associated with more reliable and stable answers, while the expected benefit of further reasoning decreases. This makes confidence a natural signal to learn: it reflects both the reliability of the current answer and the potential value of further reasoning. Learning this signal alone substantially shortens reasoning: ConfSFT reduces generated tokens by up to 25% at matched accuracy across four model families and transferring from mathematical training problems to scientific and coding tasks. Its efficiency gains are comparable to training methods that explicitly optimize for shorter reasoning. Our analyses further show that confidence prediction improves over training while generated tokens decrease and accuracy remains stable. A decomposition of reasoning traces into high-level reasoning behaviors shows that ConfSFT largely preserves the base models’ reasoning composition, rather than achieving efficiency by selectively suppressing a particular type of reasoning. Controlled ablations further show that the supervision signal matters: alternative targets, including randomly shuffled confidence labels, produce weaker or no efficiency gains. Together, these results provide evidence that learning confidence itself can reshape reasoning toward greater efficiency. 2 Related Work Training LLMs for efficient reasoning. A growing line of work directly trains reasoning models to use less computation. Several approaches use reinforcement learning objectives that explicitly favor shorter correct reasoning traces ( Liu et al., 2025 ; Arora and Zanette, 2025 ; Yi et al., 2025 ) . Others fine-tune on self-generated concise trajectories, either by selecting short reasoning paths or filtering on-policy generations for correctness and conciseness ( Munkhbat et al., 2025 ; Zhao et al., 2026 ) . Other methods explicitly train models to adapt or terminate their reasoning through truncated rollouts, selective reasoning, or confidence-guided post-training ( Chen et al., 2026a ; Huang et al., 2026 ; Jiang et al., 2026 ; Han et al., 2026 ; Qiao et al., 2025 ) . While these methods differ in how efficiency is encouraged, they all make shorter or more efficient reasoning an explicit part of the training procedure. In contrast, our approach supervises only confidence, without any explicit training signal that prefers shorter reasoning. Inference-time adaptive reasoning and early stopping. A complementary line of work reduces reasoning cost by adapting computation at inference time. Several methods use confidence or uncertainty signals to determine when further reasoning is unnecessary ( Yang et al., 2025 ; Hosseini et al., 2026 ; Wang et al., 2025 ; Kim et al., 2026 ; Yong et al., 2025 ; Fu et al., 2025 ) . Others rely on the convergence or stability of intermediate answers to terminate reasoning ( Liu and Wang, 2025 ; Mao et al., 2026 ) , while Zhang et al. (2025) train probes on internal representations to detect intermediate answer correctness and enable early exit. Although these approaches use different signals, they all explicitly use an inference-time criterion to control how much reasoning is performed. In contrast, ConfSFT does not use confidence or any other stopping criterion during inference. Confidence and metacognitive learning. A growing body of work incorporates confidence and metacognitive signals into training. Jang et al. (2025) show that fine-tuning on verbalized confidence labels alone can induce changes in reasoning behavior, including emergent self-verification. Other work explicitly trains models to improve their confidence calibration, for example by optimizing verbalized confidence calibration through reinforcement learning ( Bani-Harouni et al., 2026 ; Damani et al., 2026 ) , shaping confidence dynamics to avoid premature commitment ( Gai et al., 2026 ) , or using metacognitive self-evaluation as feedback during reinforcement learning ( Liu et al., 2026 ) . Confidence has also been incorporated into reward and alignment objectives to improve reasoning quality and reliability ( He et al., 2025a ; Chen et al., 2026b ) . ConfSFT differs in both the source and role of the confidence signal. We derive confidence targets directly from the model’s own token probabilities at intermediate reasoning states and train only to predict these targets, without gold correctness labels or external feedback. Rather than optimizing efficiency, correctness, or a desired confidence trajectory, we study whether this self-supervised confidence learning alone can induce more efficient reasoning. 3 Confidence and the Value of Further Reasoning Figure 2: (a) Trial-answer correctness and (b) agreement with the final answer increase with confidence. (c) The model often continues generating after first reaching c t ≥ 0.95 c_{t}\geq 0.95 . (d) The expected accuracy gain from further reasoning decreases with confidence. What do we mean by confidence? At an intermediate point in reasoning, we probe the model for its current answer. Given the reasoning generated so far, we append a fixed answer-elicitation prompt (e.g., The final answer is \boxed{ ) and greedily generate the model’s current trial answer, as illustrated in Figure 3 a (Step 2). Let t t index such an intermediate state and a t a_{t} denote its trial answer. We define the corresponding confidence c t c_{t} as the geometric mean of the token probabilities assigned to a t a_{t} , i.e., its length-normalized likelihood. Importantly, this score is computed entirely from the model’s own distribution and does not require knowing whether the trial answer is correct. The exact answer-elicitation prompts are provided in Appendix A.3 . We do not interpret this score as a calibrated probability of correctness: a confidence of 0.8 0.8 , for example, does not necessarily imply that the answer is correct 80 % 80% of the time. Instead, we ask whether it provides a meaningful signal about the model’s intermediate reasoning state and the value of continuing to reason. Confidence is informative about the reasoning state. We evaluate confidence at intermediate reasoning states immediately preceding each Wait in 480 Nemotron traces on AIME 2024 (16 samples per problem) and group the resulting intermediate states into equal-sized confidence quantiles, with roughly 1,700 states per bin. As shown in Figure 2 , higher confidence is strongly associated with a higher probability that the trial answer is correct. The relationship is even stronger with answer stability: as confidence increases, the trial answer becomes increasingly likely to already match the model’s eventual final answer. We observe the same trends for Gemma in Appendix D . Confidence tracks the value of further reasoning. Models nevertheless often continue reasoning after reaching a high-confidence state. After first reaching c t ≥ 0.95 c_{t}\geq 0.95 , Nemotron generates a median of 1,766 additional tokens, with many trajectories continuing for more than 10K tokens (Figure 2 ). To quantify whether this additional computation is useful, we define the accuracy gain from continuing at an intermediate state t t as U t = 𝟙 [ a final is correct ] − 𝟙 [ a t is correct ] , U_{t}=\mathbbm{1}[a_{\mathrm{final}}\ \text{is correct}]-\mathbbm{1}[a_{t}\ \text{is correct}], (1) where a t a_{t} is the trial answer at t t and a final a_{\mathrm{final}} is the answer after completing the reasoning trajectory. Thus, U t > 0 U_{t}>0 when further reasoning corrects the current answer and U t < 0 U_{t}<0 when it changes a correct intermediate answer into an incorrect final answer. As shown in Figure 2 , the expected gain from continued reasoning decreases sharply with confidence: further reasoning is most beneficial at low-confidence states and provides little average improvement once confidence is high. Together, these observations suggest that confidence provides a self-supervised signal about both the model’s current reasoning state and the value of additional computation. Prior work has used the same confidence signals to explicitly control reasoning, either through inference-time early stopping ( Yang et al., 2025 ; Hosseini et al., 2026 ) or by truncating rollouts during training ( Chen et al., 2026a ) . We instead ask a different question: what happens if the model is simply trained to predict this signal at intermediate reasoning states? In the following sections, we show that confidence supervision alone can substantially reduce generated tokens, without truncating trajectories or explicitly encouraging shorter reasoning. 4 Method Motivated by the observations in Section 3 , we introduce ConfSFT , a self-supervised procedure for fine-tuning reasoning models with confidence supervision at intermediate reasoning states. As illustrated in Figure 3 , ConfSFT repeatedly generates reasoning trajectories from the current policy, constructs confidence labels from intermediate states, and fine-tunes the model to predict these labels. We describe each step below. Generating reasoning rollouts. We begin with a small subset 𝒟 train \mathcal{D}{\mathrm{train}} of the training problems. For each problem x ∈ 𝒟 train x\in\mathcal{D}{\mathrm{train}} , we use the current policy π θ \pi_{\theta} to generate one or more reasoning trajectories τ = ( y 1 , … , y T ) \tau=(y_{1},\ldots,y_{T}) . These are standard reasoning rollouts, generated without confidence elicitation or any instruction to reason efficiently. The prompts are provided in Appendix A.1 . Constructing confidence labels. For each rollout, we identify a set of intermediate decision points using a fixed textual marker. By default, we use occurrences of Wait and take the reasoning prefix immediately preceding each occurrence as an intermediate reasoning state. Let t t index such a decision point and y < t y_{<t} denote the corresponding reasoning prefix. When a trajectory contains many decision points, we retain at most M M , selected approximately uniformly across the trajectory, to prevent a single rollout from contributing disproportionately to the training set. To construct a confidence target for y < t y_{<t} , we append a fixed answer-elicitation prompt q q and greedily generate a trial answer a = ( a 1 , … , a n ) a=(a_{1},\ldots,a_{n}) . We define its confidence as the geometric mean of the token probabilities assigned to the generated answer: c t = exp ( 1 n ∑ i = 1 n log π θ ( a i ∣ x , y < t , q , a < i ) ) . c_{t}=\exp\left(\frac{1}{n}\sum_{i=1}^{n}\log\pi_{\theta}\left(a_{i}\mid x,y_{<t},q,a_{<i}\right)\right). (2) This corresponds to the length-normalized likelihood of the trial answer. Importantly, c t c_{t} is computed entirely from the model’s own distribution and does not require the gold answer. We represent confidence as a textual percentage so that it can be learned using the model’s standard next-token prediction objective. Specifically, we quantize c t ∈ [ 0 , 1 ] c_{t}\in[0,1] onto a grid of K = 50 K=50 percentage levels, corresponding to 2 % 2% increments, and round each score up to the smallest grid point greater than or equal to it. The resulting textual percentage ℓ t \ell_{t} serves as the supervision target for decision point t t . Confidence-supervised fine-tuning. For each decision point, we construct one training example by concatenating the problem prompt, the corresponding reasoning prefix y < t y_{<t} , a fixed natural-language priming prefix p p , and the quantized confidence label ℓ t \ell_{t} : s = prompt ( x ) ⊕ y < t ⊕ p ⏟ context, masked ⊕ ℓ t ⏟ supervised s=\underbrace{\mathrm{prompt}(x)\oplus y_{<t}\oplus p}{\text{context, masked}}\oplus\underbrace{\ell{t}}{\text{supervised}} (3) where ⊕ \oplus denotes concatenation. We use the fixed priming prefix: From 0% (very low) to 100% (very high), my confidence in the answer so far is Let S ( s ) S(s) denote the token positions corresponding to the confidence label ℓ t \ell{t} . We optimize the standard next-token cross-entropy only over these positions: ℒ ( θ ) = − 𝔼 ( x , y , t ) [ ∑ j ∈ S ( s ) log π θ ( s j ∣ s < j ) ] . \mathcal{L}(\theta)=-\mathbb{E}{(x,y,t)}\left[\sum{j\in S(s)}\log\pi_{\theta}\left(s_{j}\mid s_{<j}\right)\right]. (4) All other tokens, including the problem prompt, reasoning prefix, and priming prefix, are masked from the loss. Thus, the model is trained only to predict the confidence label conditioned on an intermediate reasoning state; it is not trained to reproduce the sampled reasoning trajectory or to generate confidence text within its chain of thought. The objective contains no term for reasoning length, stopping, or efficiency. Iterative on-policy training. Because both reasoning traces and confidence targets are policy-generated, we refresh them as the policy changes. At round r r , we sample rollouts from a fresh subset of training problems using π θ r \pi_{\theta_{r}} , construct confidence labels at decision points, and fine-tune on the resulting examples to obtain π θ r + 1 \pi_{\theta_{r+1}} (Figure 3 a). Later rounds generally yield more efficient policies. We evaluate the policy after each round on a fixed held-out validation set. Inference. Inference is unchanged from the base model. The fine-tuned model uses the same decoding procedure, with no confidence instruction, answer-elicitation cue, priming prefix, verifier, or early-exit mechanism. The fine-tuned models also do not spontaneously emit confidence values or the priming prefix within their reasoning traces. Thus, any change in reasoning length arises from the fine-tuned policy itself rather than an inference-time intervention. Figure 3 b illustrates this behavior on an AIME problem, where the fine-tuned policy reaches the same correct answer with substantially shorter reasoning. (a) Training Generate rollouts Construct confidence labels Confidence-supervised fine-tuning 1 2 3 repeat 1 Problem Reasoning model ∑ 2 a − 1 = 2024 \textstyle\sum 2^{a-1}=2024 Wait 2024 = 11111101000 2 2024=11111101000_{2} Wait sum = 55 =55 2 ∑ 2 a − 1 = 2024 \textstyle\sum 2^{a-1}=2024 Final answer is: reasoning so far answer-elicitation prompt Reasoning model 5 5 one token at a time p 1 = .80 p_{1}{=}.80 p 2 = .65 p_{2}{=}.65 geometric mean 72% 3 Problem ∑ 2 a − 1 = 2024 \textstyle\sum 2^{a-1}=2024 From 0% (very low) to 100% (very high), my confidence in the answer so far is 72% masked loss (b) Inference Problem. Exactly 2024 2024 finite nonempty sets B B of positive integers have their maximum in A A . Find ∑ a ∈ A a \textstyle\sum_{a\in A}a . Base ≈ \approx 14K tokens Each a a gives 2 a − 1 2^{a-1} sets, so ∑ 2 a − 1 = 2024 \sum 2^{a-1}=2024 . Wait write 2024 2024 in binary: 11111101000 2 11111101000_{2} . Wait recheck 1024 + 512 + … + 8 = 2024 1024{+}512{+}\dots{+}8=2024 . …the same fact re-derived six times … Final answer: 55 Ours ≈ \approx 3K tokens Each a a gives 2 a − 1 2^{a-1} sets, so ∑ 2 a − 1 = 2024 \sum 2^{a-1}=2024 . 2024 = 11111101000 2 ⇒ 2024=11111101000_{2}\Rightarrow exponents { 10 , 9 , 8 , 7 , 6 , 5 , 3 } {10,9,8,7,6,5,3} . So A = { 11 , 10 , 9 , 8 , 7 , 6 , 4 } A={11,10,9,8,7,6,4} . Wait check 11 + 10 + 9 + 8 + 7 + 6 + 4 = 55 11{+}10{+}9{+}8{+}7{+}6{+}4=55 . Final answer: 55 Figure 3: Overview of ConfSFT. (a) The policy generates reasoning rollouts, constructs self-supervised confidence labels at intermediate states, and is fine-tuned only on the confidence targets. (b) At inference, the fine-tuned policy uses standard generation but can produce substantially shorter reasoning while reaching the same correct answer. 5 Experiments 5.1 Setup Models. We evaluate ConfSFT on four reasoning model families of different scales: Gemma-4-E2B ( Team, 2026 ) , Qwen3-4B ( Qwen Team, 2025 ) , Nemotron-Nano-8B ( Bercovich et al., 2025 ) , and GPT-OSS-20B ( OpenAI, 2025 ) . Training Data. We train on problems from AIME2000–2023 ( Veeraboina, 2023 ) . AIME2024 is held out for validation; no test benchmark is used during training or model selection. Evaluation. We evaluate on AIME2025 ( Dekoninck et al., 2026 ) , GSM8K ( Cobbe et al., 2021 ) , GPQA-Diamond ( Rein et al., 2024 ) , HumanEval ( Chen et al., 2021 ) , and LiveCodeBench ( Jain et al., 2024 ) . We sample 16 completions per problem and report accuracy and average generated tokens across them. The full evaluation protocol is provided in Appendix C . Training Details. We apply the ConfSFT procedure described in Section 4 . We partition the training problems into eight groups, allowing up to eight training rounds, and sample eight rollouts per training problem and 16 per validation problem. Full training hyperparameters and implementation details are provided in Appendix A . Baselines. We compare against the corresponding base model and DEER ( Yang et al., 2025 ) , an inference-time early-stopping method that uses confidence to terminate reasoning. As training-time baselines, we evaluate A&Z ( Arora and Zanette, 2025 ) , which explicitly incorporates response length into the RL objective, on Nemotron and Gemma, and On-Policy SFT ( Zhao et al., 2026 ) on Nemotron, Gemma, and Qwen. Full baseline training details and more experiments are provided in Appendix F . Metrics. We report accuracy ( Acc. ), average number of generated tokens ( Avg. Tokens ), and token reduction ( Token Red. ), defined as the percentage reduction in average generated tokens relative to the corresponding base model. We also report pass@8 ( P@8 ) in Appendix C 5.2 Main Results Table 1: Main results. ConfSFT reduces generated tokens while maintaining base-model accuracy across model families and task domains. Method Math Science Coding Average AIME2025 GSM8K GPQA-Diamond LiveCodeBench HumanEval Acc ↑ \uparrow Avg. Tokens ↓ \downarrow Acc ↑ \uparrow Avg. Tokens ↓ \downarrow Acc ↑ \uparrow Avg. Tokens ↓ \downarrow Acc ↑ \uparrow Avg. Tokens ↓ \downarrow Acc ↑ \uparrow Avg. Tokens ↓ \downarrow Acc ↑ \uparrow Token Red. Nemotron-Nano-8B Base 42.5 11,325 91.8 1,192 52.5 7,338 44.2 11,761 89.7 3,600 64.1 0.0% DEER ( Yang et al., 2025 ) 39.8 9,076 (-19.9%) 89.2 765 (-35.8%) 49.8 5,712 (-22.2%) 17.4 1,731 (-85.3%) 47.1 511 (-85.8%) 48.7 -49.8% On-Policy SFT ( Zhao et al., 2026 ) 39.0 10,487 (-7.4%) 91.1 999 (-16.2%) 50.3 6,146 (-16.2%) 41.8 11,364 (-3.4%) 89.3 3,174 (-11.8%) 62.3 -11.0% A&Z ( Arora and Zanette, 2025 ) 41.5 9,644 (-14.8%) 91.0 1,218 (+2.2%) 50.9 6,738 (-8.2%) 44.5 11,346 (-3.5%) 88.4 3,489 (-3.1%) 63.3 -5.5% ConfSFT (ours) 45.6 10,153 (-10.4%) 91.6 1,010 (-15.3%) 52.8 6,278 (-14.4%) 45.1 11,151 (-5.2%) 90.8 3,225 (-10.4%) 65.2 -11.1% Gemma-4-E2B Base 35.0 7,284 91.2 956 42.6 3,294 42.1 7,543 93.2 2,210 60.8 0.0% DEER ( Yang et al., 2025 ) 24.2 4,924 (-32.4%) 90.4 912 (-4.6%) 36.5 2,432 (-26.2%) 31.3 5,333 (-29.3%) 82.9 1,905 (-13.8%) 53.1 -21.3% On-Policy SFT ( Zhao et al., 2026 ) 27.3 5,601 (-23.1%) 90.9 1,049 (+9.7%) 42.1 3,587 (+8.9%) 41.0 7,711 (+2.2%) 93.2 2,318 (+4.9%) 58.9 +0.5% A&Z ( Arora and Zanette, 2025 ) 35.4 7,018 (-3.7%) 91.3 963 (+0.7%) 41.5 3,224 (-2.1%) 41.2 7,291 (-3.3%) 93.6 2,200 (-0.5%) 60.6 -1.8% ConfSFT (ours) 32.1 5,987 (-17.8%) 91.4 938 (-1.9%) 40.8 2,968 (-9.9%) 42.4 6,403 (-15.1%) 93.8 2,059 (-6.8%) 60.1 -10.3% Qwen3-4B Base 59.2 13,268 94.8 2,292 53.0 9,075 46.8 14,954 93.3 3,495 69.4 0.0% DEER ( Yang et al., 2025 ) 40.4 9,086 (-31.5%) 84.8 580 (-74.7%) 47.1 3,621 (-60.1%) 25.8 10,374 (-30.6%) 50.8 756 (-78.4%) 49.8 -55.0% On-Policy SFT ( Zhao et al., 2026 ) 58.8 11,405 (-14.0%) 95.1 1,414 (-38.3%) 53.9 7,694 (-15.2%) 45.9 14,518 (-2.9%) 93.1 2,944 (-15.7%) 69.3 -17.2% ConfSFT (ours) 58.8 11,365 (-14.3%) 94.5 1,710 (-25.4%) 55.8 6,838 (-24.6%) 45.0 12,122 (-18.9%) 93.3 3,054 (-12.6%) 69.5 -19.2% gpt-oss-20b Base 52.7 5,802 94.0 482 61.0 3,732 53.5 5,034 95.9 917 71.4 0.0% DEER ( Yang et al., 2025 ) 52.7 5,807 (+0.1%) 94.0 482 (+0.1%) 61.0 3,753 (+0.6%) 51.1 4,824 (-4.2%) 90.6 832 (-9.2%) 69.9 -2.5% ConfSFT (ours) 53.3 5,235 (-9.8%) 94.1 437 (-9.2%) 61.0 3,112 (-16.6%) 50.7 4,393 (-12.7%) 95.0 850 (-7.2%) 70.8 -11.1% Table 1 shows that ConfSFT consistently reduces generated tokens while maintaining base-model accuracy across all four model families. ConfSFT reduces tokens by 11.1%, 10.3%, 19.2%, and 11.1% on Nemotron, Gemma, Qwen, and GPT-OSS, respectively. In Appendix C.5 , we report confidence intervals and show that the token reductions are statistically significant, while the changes in accuracy are not. Importantly, although ConfSFT is trained only on AIME problems, the efficiency gains transfer beyond mathematics to scientific reasoning and coding benchmarks. Comparison with explicit efficiency training. We compare ConfSFT with two training-time approaches: A&Z ( Arora and Zanette, 2025 ) and On-Policy SFT ( Zhao et al., 2026 ) . Across the model families where these baselines are evaluated, ConfSFT achieves comparable or larger efficiency gains despite receiving no supervision that favors shorter reasoning. For example, on Qwen, On-Policy SFT reduces tokens by 17.2% at 69.3 average accuracy, compared with 19.2% at 69.5 accuracy for ConfSFT. On Nemotron, ConfSFT similarly achieves substantial token reduction while maintaining base-model accuracy. These results show that confidence-only supervision can yield efficiency gains comparable to methods explicitly designed to shorten reasoning. Comparison with inference-time early stopping. We additionally compare against DEER ( Yang et al., 2025 ) , which uses confidence to explicitly terminate reasoning at inference time. DEER can produce larger token reductions, but its effectiveness varies considerably across tasks and can come with large accuracy losses. This is particularly pronounced on coding benchmarks: with Nemotron, LiveCodeBench accuracy drops from 44.2% to 17.4% and HumanEval accuracy from 89.7% to 47.1%. In contrast, ConfSFT reduces generated tokens without applying an inference-time stopping rule and preserves accuracy across domains. Additional baseline details are provided in Appendix F . 5.3 Confidence Learning and Efficiency Dynamics We next examine whether the model actually learns the confidence-prediction task, and how this relates to the efficiency gains during training. We measure prediction quality using confidence mean absolute error (C-MAE), the mean absolute difference between the model’s predicted confidence and the quantized self-supervised target on held-out reasoning states. Predictions are elicited using the same priming prefix as in Section 4 ; lower C-MAE indicates better prediction of the confidence targets. Figure 4 (a–c) shows the training dynamics for Nemotron on the validation set. As training progresses, C-MAE decreases substantially and average generated tokens fall, while accuracy remains approximately unchanged. Thus, the model learns to predict the confidence targets at the same time that its reasoning becomes more efficient. Figure 4 (d) further shows how the predicted confidence distribution changes. Before training, the model concentrates its predictions on only a few confidence levels. After ConfSFT, predictions span a much broader range, suggesting that the model learns to make more graded confidence judgments. We observe similar confidence-learning dynamics for Gemma in Appendix D . Figure 4: Training Dynamics. (a) Accuracy remain stable. (b) Generated tokens decrease. (c) Confidence predictions improved. (d) Confidence predictions become more graded. Figure 5: Different efficiency methods reshape reasoning in different ways. Following the Schoenfeld-style episode taxonomy, we compare each efficient policy with its corresponding base policy. Left: change in reasoning-token share for each episode, in percentage points. Right: total variation distance between the before- and after-training episode distributions. 5.4 Reasoning Composition under Efficiency Training Prior work has used Schoenfeld’s Episode Theory to decompose model reasoning into different cognitive steps ( Li et al., 2025 ; Li et al., 2026 ) . In particular, ThinkARM ( Li et al., 2026 ) annotates reasoning at the sentence level using eight categories— Read , Analyze , Plan , Implement , Explore , Verify , Monitor , and Answer —and shows that efficiency methods can alter reasoning in qualitatively different ways, selectively changing the amount of computation allocated to different episodes rather than simply shortening all reasoning uniformly. Following this framework, we annotate reasoning traces before and after efficiency training and compare the fraction of reasoning tokens assigned to each episode. We compare ConfSFT with several explicit efficiency-training methods, including L1-Max ( Aggarwal and Welleck, 2025 ) , ThinkPrune ( Hou et al., 2025 ) , A&Z ( Arora and Zanette, 2025 ) , and On-Po The alibaba story also surfaces in Qwen Developers Open-Source Local-First Search Layer..., adding another angle. as detailed in the full paper on Arxiv The alibaba story also surfaces in Alibaba Plans AI Model With Up..., adding another angle.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!