Back to AI Research

AI Research

Probability is Not Enough: Exploring and Counting D... | AI Research

Key Takeaways

  • What the paper is about As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidenc...
  • As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers.
  • Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear.
  • Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence.
  • We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding.
Paper AbstractExpand

As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding. DTC identifies these divergent tokens using the Jensen-Shannon divergence between next-token distributions evaluated along the same reasoning trajectory. We find that their count is almost negatively associated with answer accuracy, thereby serving as a simple yet effective signal for uncertainty quantification. DTC supports both white-box and black-box evaluation using auxiliary models, without explicit training and affecting the generation process. Experiments across multiple model families and six mathematical benchmarks demonstrate improved calibration over probability-based and verbalized baselines. Under white-box evaluation, the count-only estimator achieves an average expected calibration error of 13.0%, compared with 32.7%-42.4% for standard full-sequence confidence methods. In black-box settings, it also improves calibration over the original verbalized scores. For example, mean expected calibration error falls from 32.1%-40.2% to 13.7%-16.3% on DeepSeek-V3.2. These findings provide new insights for improving reasoning uncertainty quantification in large language models. The code is released at this https URL .

What the paper is about

As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding. DTC identifies these divergent tokens using the Jensen-Shannon divergence between next-token distributions evaluated along the same reasoning trajectory. We find that their count is almost negatively associated with answer accuracy, thereby serving as a simple yet effective signal for uncertainty quantification. DTC supports both white-box and black-box evaluation using auxiliary models, without explicit training and affecting the generation process. Experiments across multiple model families and six mathematical benchmarks demonstrate improved calibration over probability-based and verbalized baselines. Under white-box evaluation, the count-only estimator achieves an average expected calibration error of 13.0%, compared with 32.7%-42.4% for standard full-sequence confidence methods. In black-box settings, it also improves calibration over the original verbalized scores. For example, mean expected calibration error falls from 32.1%-40.2% to 13.7%-16.3% on DeepSeek-V3.2. These findings provide new insights for improving reasoning uncertainty quantification in large language models. The code is released at this https URL .

What it covers

Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs Feiyang Li Shengjing Liu Qi Zhan Sijie Cheng Affiliation: College of Computer Science and Software Engineering, Shenzhen University RayNeo.AI Weiqing Wang Hongwen Chen Yuxuan Yang Wen Wang Affiliation: College of Computer Science and Software Engineering, Shenzhen University RayNeo.AI Affiliation: Tsinghua University Behavioral and Spatial AI Lab, Peking University & Tongji University [email protected] [email protected] Yile Wang Hui Huang Abstract As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertainty of their answers. Current methods for estimating the confidence of large language models are generally based on probabilities of selected key tokens, but the underlying mechanism remains unclear. Our pilot study finds that replacing selected token probabilities with coarse substitutes can also improve calibration, motivating us to further explore effective signals of model confidence. We introduce Divergent Token Confidence (DTC), a framework that estimates confidence by counting tokens at which two models strongly disagree during decoding. DTC identifies these divergent tokens using the Jensen–Shannon divergence between next-token distributions evaluated along the same reasoning trajectory. We find that their count is almost negatively associated with answer accuracy, thereby can serve as a simple yet effective signal for uncertainty quantification. DTC supports both white-box and black-box evaluation using auxiliary models, without explicit training and affecting the generation process. Experiments across multiple model families and six mathematical benchmarks demonstrate improved calibration over the considered probability-based and verbalized-based baselines. Under white-box evaluation, the count-only estimator achieves an average expected calibration error of 13.0%, compared with 32.7%–42.4% for standard full-sequence confidence methods. In black-box settings, it also improves calibration over the original verbalized scores. For example, mean expected calibration error falls from 32.1%–40.2% to 13.7%–16.3% on DeepSeek-V3.2. These findings provide new insights for improving reasoning uncertainty quantification in large language models. The code is released at https://github.com/szu-tera/DTC.git. 1 1 footnotetext: Equal contribution. 2 2 footnotetext: Corresponding author. 1 Introduction Large language models have shown strong capability to solve complex tasks through Chain-of-Thought (CoT) reasoning ( Wei et al., 2022 ; OpenAI, 2024 ; Guo et al., 2025 ) . However, a reasoning trajectory that yields a correct answer may still contain unreliable intermediate decisions, while one that yields an incorrect answer may still be logically reliable ( Wei et al., 2022 ; Bao et al., 2025 ; Zhou et al., 2026 ) . Therefore, estimating the confidence of a model’s reasoning trajectory is becoming increasingly important. First, a mismatch between model confidence and its answers can amplify the model’s unreliability, thereby limiting its deployment in safety-critical scenarios ( Clusmann et al., 2023 ) . Second, understanding reasoning confidence can help us determine whether we need to rely on the outputs of models. When a reasoning path is identified as unreliable, the model can abstain from answering ( Madhusudhan et al., 2025 ) , or it can be deferred to human review ( Devic et al., 2025 ) . Finally, it has been shown that confidence itself can also be leveraged to improve the model’s own reasoning performance ( Fu et al., 2026 ) . We refer to this confidence estimation task as Reasoning Uncertainty Quantification : assigning a confidence score to each query and its reasoning path to reflect the reliability of the path and its final answer ( Liu et al., 2025 ; Zhang & Zhang, 2025 ) . Existing methods typically aggregate token probabilities over the entire trajectory, but the resulting scores can be overconfident ( Orgad et al., 2025 ; Li et al., 2025 ; Zhang & Zhang, 2025 ) . To mitigate this overconfidence, Li et al. (2025) propose Uncertainty Quantification with Attention Chain (UQAC), which selects answer-related CoT tokens and multiplies their probabilities to obtain a path-level confidence score. This design suggests that the probabilities of answer-related tokens provide useful information for estimating confidence in answer correctness. However, its product score depends jointly on these probabilities and the selected-token count, leaving it unclear whether the calibration gains come from the selected probabilities or the count. To test whether the selected probabilities are necessary for these calibration gains, we keep UQAC’s selected token set fixed and replace the selected probabilities with either the trajectory’s mean token probability or a constant shared across trajectories from the same model on a given dataset. Both replacements generally improve calibration in our case study, suggesting that the selected tokens’ specific probabilities may not be necessary for these gains. In particular, replacing the probabilities with a shared constant makes the score depend only on the number of selected tokens, motivating us to examine token count as a calibration signal ( Section 3.3 ). To obtain the token count signal that reflects reasoning path reliability, we draw on inter-model disagreement as an uncertainty signal ( Kruse et al., 2025 ; Sun et al., 2024 ) . We hypothesize that unreliable reasoning paths contain more tokens at which models strongly disagree . We measure token-level disagreement using the Jensen–Shannon divergence (JSD) between two models’ next-token distributions and define tokens whose divergence exceeds a threshold θ \theta as divergent tokens. Their count summarizes the frequency of strong disagreement along the reasoning path and is negatively associated with path accuracy in our experiments ( Figure 1 ). We therefore introduce Divergent Token Confidence (DTC) , which converts divergent-token count into path-level confidence. DTC provides two estimators: DTC lin \mathrm{DTC}{\mathrm{lin}} maps the count directly to confidence, while DTC prod \mathrm{DTC}{\mathrm{prod}} combines the count with a full-sequence probability score. The same procedure can be applied in white-box settings using the generator’s information and in black-box settings using auxiliary models on the generated trajectory, without requiring access to the generator’s logits. Across multiple model families and mathematical benchmarks, DTC improves calibration over full-sequence likelihood-based, verbalized-confidence, and UQAC baselines ( Sections 5.1 and 5.2 ). DTC also reduces overconfidence in verbalized scores on existing trajectories ( Section 6 ). Figure 1: Divergent-token selection and its relationship to path accuracy. (a) A generator produces a reasoning trajectory. (b) A token is selected when the JSD between two models’ next-token distributions exceeds θ \theta . (c) Path accuracy decreases as the number of divergent tokens increases under single- and dual-auxiliary probing. 2 Related Work Probability-based confidence and verbalized confidence. In white-box uncertainty quantification, token probabilities or predictive uncertainty from the generating model are commonly used to construct response-level confidence scores. Sequence-level uncertainty estimators typically aggregate token-level signals over a generated response ( Malinin & Gales, 2021 ) . Relevance-aware methods account for differences in how tokens contribute to meaning or the final answer, weighting these signals by token relevance, sentence relevance, or contextual information ( Duan et al., 2024 ; Bakman et al., 2024 ; Lin et al., 2024 ) . For responses that include CoT reasoning, CoT-UQ and UQAC further use the relevance of tokens in the reasoning chain to the final answer to estimate response-level uncertainty ( Zhang & Zhang, 2025 ; Li et al., 2025 ) . In black-box settings, verbalized confidence is a common approach to confidence estimation ( Wang & Zhang, 2026 ) . Such methods either elicit confidence through additional prompts after a response has been generated or ask the model to report confidence or a distribution over candidate answers alongside its CoT and final answer ( Tian et al., 2023 ; Wang et al., 2025 ) . The latter couples reasoning with confidence elicitation: the choice of prompt may alter the CoT trajectory and lead to lower accuracy or overconfidence in mathematical reasoning ( Wang et al., 2025 ) . Our work reveals their limitations and investigates the novel divergent-token count as a confidence signal in both white-box and black-box settings. Cross-model signals for confidence estimation. Another line of work uses multiple models, base models, or model perturbations to estimate the reliability of generated responses. MUSE and CrossCheckGPT use multi-model consensus or cross-system consistency to improve confidence estimation ( Kruse et al., 2025 ; Sun et al., 2024 ) . BaseCal uses signals from base models to calibrate the confidence of post-trained models ( Tan et al., 2026 ) . More closely related to our work are methods that estimate reasoning uncertainty from token entropy. Some aggregate the generating model’s token entropy along a CoT trajectory; others use randomly perturbed models or additional models and aggregate their token entropy along the same trajectory to estimate uncertainty over the reasoning trajectory ( Zhang et al., 2026 ; Gorbett & Jana, 2026 ) . These studies primarily focus on response-level uncertainty or its ranking performance, whereas our focus is on a bounded score that can be directly interpreted as confidence and used for calibration. 3 Preliminaries and Pilot Study 3.1 Reasoning Uncertainty and Calibration The prediction uncertainty (or query uncertainty) U ⁡ ( x ) U(x) characterizes uncertainty over answers to a query x x before conditioning on a particular realized reasoning path, and is often estimated by sampling multiple answers to the same query ( Kuhn et al., 2023 ) . In this work, we study reasoning uncertainty U ⁡ ( x , r ) U(x,r) , the uncertainty in final-answer reliability conditioned on both the query x x and the realized reasoning process r r ( Liu et al., 2025 ; Zhang & Zhang, 2025 ) . Given the query x x , a generating model 𝒢 {\mathcal{G}} produces a trajectory τ = ( τ 1 , … , τ T ) \tau=(\tau_{1},\ldots,\tau_{T}) of T T tokens, consisting of a CoT path r r and a final answer y y ; a reference answer or evaluator provides the binary correctness label z ∈ { 0 , 1 } z\in{0,1} . We operationalize U ⁡ ( x , r ) U(x,r) through a calibrated confidence C ⁡ ( y , x , r ) ∈ [ 0 , 1 ] C(y,x,r)\in[0,1] that estimates P ⁡ ( z = 1 ∣ x , r , y ) P(z{=}1\mid x,r,y) , with higher confidence indicating lower reasoning uncertainty. Uncertainty is perfectly calibrated if P ⁡ ( z = 1 ∣ y , x , r ) = q , s.t. ​ C ​ ( y , x , r ) = q . P\bigl(z{=}1\mid y,x,r\bigr)=q,\text{ s.t. }C(y,x,r)=q. (1) Expected calibration error (ECE; Guo et al., 2017 ) is used to quantify the degree of miscalibration: we partitions N N scored trajectories into M M equal-width confidence bins { B 1 , … , B M } {B_{1},\ldots,B_{M}} and then calculates the absolute gap between average accuracy acc ⁡ ( B m ) \mathrm{acc}(B_{m}) and confidence conf ⁡ ( B m ) \mathrm{conf}(B_{m}) : ECE = ∑ m = 1 M | B m | N ​ | acc ⁡ ( B m ) − conf ⁡ ( B m ) | . \mathrm{ECE};=;\sum_{m=1}^{M}\frac{|B_{m}|}{N},\bigl|\mathrm{acc}(B_{m})-\mathrm{conf}(B_{m})\bigr|. (2) Lower ECE indicates better calibration. Unless noted, we report ECE with M = 20 M{=}20 . 3.2 Sequence Confidence from Token Probabilities Standard full-sequence confidence scores aggregate token probabilities over the entire trajectory. One such score is the length-normalized sequence likelihood (NSL; Malinin & Gales, 2021 ). Given the generator’s token probabilities p t = P 𝒢 ​ ( τ t ∣ x , τ θ \mathrm{JSD}(P_{t},Q_{t})>\theta and set of divergent token positions along a reasoning path is: T div ​ ( θ ) = { t ∈ { 1 , … , T } : JSD ⁡ ( P t , Q t ) > θ } . T_{\rm div}(\theta);=;\bigl{,t\in{1,\dots,T}:\mathrm{JSD}(P_{t},Q_{t})>\theta,\bigr}. (7) We use JSD for its symmetry and boundedness and see Section C.4 for discussion on alternative measures of token-level disagreement. Based on the definition, we can first measure the confidence of generator 𝒢 {\mathcal{G}} with an auxiliary model 𝒜 {\mathcal{A}} in the single-auxiliary setting, which applies when 𝒢 {\mathcal{G}} is a white-box model. When the target model is a black-box model and we cannot access its token distribution, we assume that the relationship between the number of divergent tokens and reasoning-path reliability remains reasonably stable across different model pairs , which allow us to calibrate in a dual-auxiliary setting by comparing two external auxiliary models from the same family, denoted as 𝒜 ′ {\mathcal{A}}^{\prime} and 𝒜 ′′ {\mathcal{A}}^{\prime\prime} . To verify that the number of divergent tokens | T div ​ ( θ ) | \lvert T_{\rm div}(\theta)\rvert can serve as a signal of model uncertainty, we first analyze the relationship between answer accuracy against the number of divergent tokens for different model pairs. The results are shown in Figure 1 (c). In both single- and dual-auxiliary settings, accuracy decreases nearly monotonically as | T div ​ ( θ ) | \lvert T_{\rm div}(\theta)\rvert increases, supporting that the number of divergent tokens can reveal reasoning-path reliability to some extent. 4 DTC : Divergent Token Confidence Given the divergent-token count m = | T div ​ ( θ ) | m=\lvert T_{\rm div}(\theta)\rvert defined in Equation 7 , DTC constructs two path-level confidence estimates. The primary estimator DTC lin \mathrm{DTC}{\mathrm{lin}} maps m m directly to confidence, while DTC prod \mathrm{DTC}{\mathrm{prod}} uses m m to recalibrate the standard full-sequence confidence C mean C_{\mathrm{mean}} from Section 3.2 . 𝐃𝐓𝐂 𝐥𝐢𝐧 \mathrm{DTC}{\mathrm{lin}} : Count-linear confidence. Motivated by the decrease in accuracy with increasing divergent token count ( Figure 1 (b,c)), we define DTC lin ​ ( m ) = { a − a − b n ​ m , 0 ≤ m < n , b , m ≥ n . \mathrm{DTC}{\mathrm{lin}}(m)=\begin{cases}a-\dfrac{a-b}{n},m,&0\leq m<n,\[4.0pt] b,&m\geq n.\end{cases} (8) Confidence starts at a a , decreases linearly with m m , and reaches a floor of b b at m = n m=n . We use a = 0.95 a=0.95 , b = 0.05 b=0.05 , and n = 10 n=10 in all main experiments; sensitivity to n n is examined in Section C.5 . Once the divergent tokens are selected, this mapping depends only on their count. 𝐃𝐓𝐂 𝐩𝐫𝐨𝐝 \mathrm{DTC}{\mathrm{prod}} : Trajectory-mean product confidence. Section 3.3 shows that replacing selected-token probabilities with the trajectory mean C mean C{\mathrm{mean}} and taking their product reduces overconfidence. We therefore combine C mean C_{\mathrm{mean}} with the uncertainty-informed count m m for confidence estimation: DTC prod ​ ( m ) = C mean m + k . \mathrm{DTC}{\mathrm{prod}}(m)=C{\mathrm{mean}}^{,m+k}. (9) For C mean C_{\mathrm{mean}} , we use probabilities from the generator in white-box settings ( Section 5.1 ) and the larger auxiliary on the same frozen trajectory in black-box settings ( Section 5.2 ). We use k = 4 k=4 unless otherwise specified, avoiding a score of one solely due to a zero count and sensitivity to k k is examined in Section C.7 . When C mean ∈ ( 0 , 1 ) C_{\mathrm{mean}}\in(0,1) , a larger m m yields a lower score. Compared with DTC lin \mathrm{DTC}{\mathrm{lin}} , DTC prod \mathrm{DTC}{\mathrm{prod}} further uses the trajectory-level probability C mean C_{\mathrm{mean}} , giving finer-grained confidence to trajectories that share the same m m . The uncertainty-informed count m m can also be combined with other overconfident scores to improve calibration ( Section 6 ). Table 1: Uncertainty quantification performance (ECE) in white-box settings. In each row, the best result is in bold and the second-best one is underlined (excluding PRM). †: UQAC variant by ours. Reasoning Models Acc 𝑪 𝐍𝐒𝐋 C_{\mathrm{NSL}} 𝑪 𝐦𝐞𝐚𝐧 C_{\mathrm{mean}} Entropy Conf. BaseCal UQAC Verb. DTC (Ours) attn mean † prod lin MATH-500 Qwen2.5 - 7B 76.1 43.6 45.3 32.8 43.7 26.8 12.7 44.2 20.0 13.4 Qwen2.5 - 14B 80.0 43.0 44.9 38.9 43.0 26.8 14.2 34.7 16.2 0 9.9 Qwen2.5 - 32B 82.4 43.7 45.4 39.9 43.8 26.1 16.2 41.6 16.9 0 6.6 Qwen3 - 8B 83.9 38.1 41.5 31.8 36.7 37.1 37.8 47.8 0 7.9 15.6 Qwen3 - 14B 86.5 36.8 40.6 29.9 36.2 27.1 13.2 42.5 0 5.5 11.6 Qwen3 - 32B 83.7 36.1 40.2 29.4 — 38.8 38.6 37.0 0 6.1 11.0 Gemma3 - 12B 84.9 41.5 44.0 37.3 34.8 33.9 29.3 24.3 15.2 13.9 Gemma3 - 27B 89.2 42.3 44.5 38.4 36.0 16.0 0 7.0 33.4 15.0 0 7.1 AMC23 Qwen2.5 - 7B 53.6 43.6 45.3 31.7 44.0 24.7 0 8.2 42.5 15.7 0 6.7 Qwen2.5 - 14B 61.2 42.7 44.7 38.5 43.1 25.3 19.7 35.1 0 9.7 0 9.5 Qwen2.5 - 32B 66.4 43.3 45.1 39.5 44.0 24.9 22.7 38.3 0 9.8 10.8 Qwen3 - 8B 68.6 36.5 40.5 29.4 36.8 49.4 38.5 48.7 0 6.8 0 9.3 Qwen3 - 14B 74.1 34.6 39.2 26.9 35.9 27.2 10.5 35.3 12.3 11.8 Qwen3 - 32B 67.5 33.7 38.5 25.4 — 44.0 37.7 32.3 15.2 0 7.5 Gemma3 - 12B 66.8 40.8 43.6 36.4 35.3 34.5 24.0 25.9 12.7 10.9 Gemma3 - 27B 76.9 41.8 44.3 37.9 36.9 19.9 12.3 26.4 12.4 12.0 AIME24 Qwen2.5 - 7B 12.6 42.8 44.7 30.3 43.3 29.5 19.0 39.5 14.7 11.5 Qwen2.5 - 14B 13.8 41.8 44.0 37.5 42.3 25.1 15.6 31.4 13.1 16.9 Qwen2.5 - 32B 16.9 42.4 44.4 37.9 43.5 23.9 16.7 34.1 10.6 20.4 Qwen3 - 8B 27.7 34.9 39.3 26.5 36.4 40.3 42.9 45.8 14.2 14.9 Qwen3 - 14B 27.1 33.2 38.1 24.4 35.7 31.4 16.2 44.1 20.6 18.1 Qwen3 - 32B 28.3 32.0 37.4 23.7 — 44.2 37.6 25.7 21.6 16.4 Gemma3 - 12B 23.9 39.8 42.8 34.7 34.6 33.2 31.8 23.0 0 9.7 12.6 Gemma3 - 27B 29.0 40.7 43.5 36.2 36.1 29.6 29.4 21.1 0 8.1 24.3 AIME25 Qwen2.5 - 7B 0 9.1 42.4 44.4 31.0 43.1 26.6 16.2 37.3 11.2 20.5 Qwen2.5 - 14B 14.7 42.4 44.4 37.8 42.9 27.5 15.7 31.6 10.8 25.4 Qwen2.5 - 32B 12.2 42.7 44.7 38.3 43.8 30.6 18.4 33.0 0 8.7 28.3 Qwen3 - 8B 22.5 34.2 38.8 25.9 36.0 37.6 43.4 46.9 12.3 0 9.5 Qwen3 - 14B 27.5 32.6 37.8 24.0 35.3 30.8 19.7 53.0 17.1 0 5.3 Qwen3 - 32B 23.7 31.7 37.2 22.5 — 36.3 37.8 25.6 19.1 0 6.9 Gemma3 - 12B 18.8 40.3 43.2 35.5 34.5 32.1 29.4 15.9 16.7 0 7.8 Gemma3 - 27B 25.3 40.6 43.4 35.9 35.4 25.1 22.7 23.1 15.6 11.1 Average 48.0 39.3 42.4 32.7 39.0 30.8 23.6 35.0 13.2 13.0 PRM (ref.) 0 9.8 0 8.8 0 8.4 0 7.4 0 6.2 0 6.3 0 7.2 0 8.4 0 6.0 0 6.0 0 8.9 11.4 12.0 11.6 10.1 13.1 22.7 23.0 23.4 29.2 29.7 28.8 29.2 30.0 25.2 24.4 27.7 29.6 30.7 30.8 27.2 26.7 18.1 5 Experiments Figure 3: Calibration plots and probability histogram for Qwen2.5-14B on MATH-500. The x x -axis shows mean confidence within 20 probability bins. The calibration curve ( blue line with μ ± σ \mu\pm\sigma ) displays actual accuracy per bin, while the gray shadow represents the probability proportion. 5.1 White-box Setting Datasets and Models. To evaluate DTC on tasks with long CoT path, we use four widely used mathematical benchmarks spanning a range of difficulty: MATH-500 ( Hendrycks et al., 2021 ; Lightman et al., 2024 ) and the contest sets AMC23, AIME24, and AIME25 ( Balunovic et al., 2025 ) . We evaluate eight LLMs from three families: Qwen2.5-{7,14,32}B-Instruct ( Qwen Team, 2024 ) , Qwen3-{8,14,32}B ( Yang et al., 2025 ) , and Gemma3-{12,27}B-IT ( Gemma Team, 2025 ) . Baselines. Given the limited prior work on calibrated reasoning uncertainty, we compare DTC against three classes of uncertainty estimators. (1) Standard full-sequence confidence : 𝑪 𝐍𝐒𝐋 C_{\mathrm{NSL}} uses length-normalized sequence likelihood ( Equation 3 ); 𝑪 𝐦𝐞𝐚𝐧 C_{\mathrm{mean}} uses mean token probability ( Section 3.2 ); and Entropy Confidence that uses length-normalized predictive entropy confidence ( Li et al., 2025 ) . (2) Refinements of full-sequence confidence : BaseCal ( Tan et al., 2026 ) averages response-token probabilities under a paired base model; 𝐔𝐐𝐀𝐂 𝐚𝐭𝐭𝐧 \mathrm{UQAC}{\mathrm{attn}} ( Li et al., 2025 ) further select attention-related tokens; and our variant 𝐔𝐐𝐀𝐂 𝐦𝐞𝐚𝐧 \mathrm{UQAC}{\mathrm{mean}} which replaces each selected-token probability with the C mean C_{\mathrm{mean}} before multiplying over the selected set. BaseCal is omitted for Qwen3-32B because no public base checkpoint is available. (3) Verbalized estimation : Xiong et al. (2024) propose Verbalized confidence through additional prompts after the response elicits the model confidence. We also include PRM (Skywork-o1-Open-PRM-Qwen-2.5-7B; He et al., 2024 ) as a supervised reference, averaging rewards over newline-delimited reasoning steps into path-level confidence. Implementations. We use the officially recommended decoding parameters for each LLM, as detailed in Section A.2 . For each model family, we use its smallest model and 100 problems drawn at random from the MATH training set ( Hendrycks et al., 2021 ) to set that family’s disagreement threshold θ \theta . We report DTC prod \mathrm{DTC}{\mathrm{prod}} and DTC lin \mathrm{DTC}{\mathrm{lin}} , with θ = 0.50 \theta=0.50 for Qwen2.5, θ = 0.60 \theta=0.60 for Qwen3, and θ = 0.85 \theta=0.85 for Gemma3. For DTC prod \mathrm{DTC}{\mathrm{prod}} , token probabilities come from the generator 𝒢 {\mathcal{G}} ( Equation 9 ). Unless noted, the auxiliary is the smaller same-family instruct model: Qwen2.5-1.5B-Instruct, Qwen3-1.7B, or Gemma3-4B-IT. Other auxiliary sizes are examined in Section 6 . Following Li et al. (2025) , we evaluate ECE and AUROC by repeated subsampling, emphasizing ECE and the calibration plots. AUROC is the area under the ROC curve and measures ranking. Tables and figures use AUC for AUROC. Further details are in Section A.3 . Main Results. Table 1 and Table 5 ( Section B.1 ) report the ECE and AUROC results, respectively. Figure 3 visualizes calibration curves on MATH-500 with Qwen2.5-14B. We find that: Full-sequence confidence retains ranking information but remains overconfident. The three standard methods ( C NSL C{\mathrm{NSL}} , C mean C_{\mathrm{mean}} and Entropy Confidence) achieve an high average AUROC of 72.7 72.7 – 73.8 73.8 , yet their ECEs reach 32.7 32.7 – 42.4 42.4 . Their calibration curves lie below the diagonal, with scores concentrated near the high-confidence. Refinements improve calibration unevenly and weaken ranking. BaseCal offers limited calibration gains, whereas the UQAC variants reduce ECE more substantially. UQAC variant UQAC mean \mathrm{UQAC}{\mathrm{mean}} by ours achieves both the lowest average ECE ( 23.6 23.6 ) and highest AUROC ( 66.7 66.7 ) among these refinements, but its ranking still trails the standard methods. DTC improves both calibration and average ranking. DTC lin \mathrm{DTC}{\mathrm{lin}} and DTC prod \mathrm{DTC}{\mathrm{prod}} reduce average ECE to 13.0 13.0 and 13.2 13.2 , with calibration curves closer to the diagonal, while attaining AUROC 75.3 75.3 and 80.8 80.8 . Thus, count alone supports effective calibration; retaining path-mean probability further improves average ranking at similar ECE. Relative to UQAC mean \mathrm{UQAC}{\mathrm{mean}} , DTC prod \mathrm{DTC}{\mathrm{prod}} lowers average ECE by 10.4 10.4 , suggesting that selecting divergent tokens captures uncertainty more effectively and improves the calibration of C mean C{\mathrm{mean}} . In harder datasets, PRM has higher ECE, while DTC is better calibrated. 5.2 Black-box Setting Datasets and Models. To better match a black-box evaluation setting, we use more challenging contest benchmarks than in the white-box experiments, namely AIME24, AIME25, HMMT25, and HMMT26 ( Balunovic et al., 2025 ) , and switch to stronger generators: Qwen3-4B-Instruct-2507, Qwen3-30B-A3B-Instruct-2507 ( Yang et al., 2025 ) , and DeepSeek-V3.2 ( DeepSeek-AI, 2025 ) . We treat each generator as a black box, using its output trajectories without accessing its log The ai search story also surfaces in Stanford AI discovery identifies natural weight..., adding another angle. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!