Back to AI Research

AI Research

Diagnosing and Improving Probabilistic Reasoning in... | AI Research

Key Takeaways

  • Diagnosing and Improving Probabilistic Reasoning in Large Language Models Large language models are increasingly asked not only to estimate what might be tru...
  • Large language models (LLMs) are increasingly proposed as decision assistants who must reason probabilistically from available evidence under explicit decision costs.
  • Using a synthetic benchmark with known ground truth, we apply the decomposition to characterize probabilistic reasoning in frontier and open-sourced models.
  • # Diagnosing and Improving Probabilistic Reasoning in Large Language Models
  • Large language models are increasingly asked not only to estimate what might be true, but also to choose actions whose consequences depend on different costs.
Paper AbstractExpand

Large language models (LLMs) are increasingly proposed as decision assistants who must reason probabilistically from available evidence under explicit decision costs. We propose a decision-theoretic framework that decomposes LLMs' decision loss into two components: forming accurate beliefs from provided evidence and translating those beliefs into actions that optimize a provided utility function. Using a synthetic benchmark with known ground truth, we apply the decomposition to characterize probabilistic reasoning in frontier and open-sourced models. We further evaluate whether RL interventions targeting beliefs, decisions, or both improve these components across three domains, whether improvements transfer across components and elicitation formats, and whether decision performance can improve without improvement in belief formation. We find that targeting one component of probabilistic reasoning redistributes decision loss, improving the target without necessarily transferring to others, and that jointly targeting belief formation and decision-making improves both but hinges on matched formats between training and evaluation.

Diagnosing and Improving Probabilistic Reasoning in Large Language Models

Large language models are increasingly asked not only to estimate what might be true, but also to choose actions whose consequences depend on different costs. This paper develops a framework for separating two abilities: forming an accurate probability estimate from evidence, and using that estimate to choose the best action under an explicit loss function. The authors use controlled tasks with known answers to diagnose where decision errors come from, then test whether reinforcement-learning interventions can improve one or both abilities.

What the paper measures

The framework focuses on binary outcomes, such as whether a condition is present or whether severe weather will occur. For each problem, the model sees evidence and is asked to report a probability—for example, a 70% chance that the outcome is positive. The benchmark’s data-generating process provides a reference posterior probability, making it possible to assess the model’s belief formation directly.
The model must also choose an action under a specified loss function. A loss function assigns different costs to mistakes: a false positive may be more or less costly than a false negative. The optimal action therefore depends on both the probability of the outcome and the relative costs. The same belief can rationally produce different actions when the loss function changes.
The authors distinguish three actions. The optimal action is the one that minimizes expected loss under the reference probability. The belief-implied action is the action that would be optimal if the model’s reported probability were used. The model’s reported action is what it actually chooses. Comparing these actions decomposes total regret into two components.
Belief loss is the cost of using the model’s reported belief rather than the true posterior. It can result from an incorrect prior, failure to extract relevant information, non-Bayesian updating, or distortions caused by how the question is elicited. Optimization residual captures the additional discrepancy between the action implied by the model’s reported belief and the action the model actually selects. This residual can even be negative if the reported action performs better than the action implied by the model’s stated probability.

How the evaluation works

The first evaluation uses a synthetic benchmark in which the correct posterior and optimal action are known by construction. Each instance contains labeled observations, a binary test case, and a decision problem. The authors vary the number of binary features, from none to five, so they can increase the difficulty of extracting the relevant statistical relationship.
Models report beliefs under a quadratic scoring rule and make decisions across nine loss functions. These loss functions vary the relative costs of false positives and false negatives, producing decision thresholds from 0.1 to 0.9. Regrets are normalized to account for different loss scales, while preserving the decomposition into belief loss and optimization residual.
The authors then test four reinforcement-learning interventions using Qwen3-8B: belief-only training, decision-only training, separate multitask training, and sequential alignment training. Belief-only training rewards accurate probability reports. Decision-only training rewards choosing the optimal action, without directly supervising the belief. Separate multitask training applies the first two forms of supervision on different examples.
Sequential alignment training asks the model to report a belief first and then choose an action using that belief and the loss function. It rewards both accurate beliefs and consistency between the reported belief and the selected action. The interventions are evaluated across synthetic inference, weather forecasting, and clinical reasoning. These settings range from structured, highly controlled evidence to natural-language descriptions and clinical notes, while retaining known or model-derived reference posteriors.

Results that stand out

The diagnostic results show that models do not have one fixed pattern of probabilistic failure. In simple settings with no features or one feature, several frontier models exhibit perfect belief-decision alignment under the tested conditions. As the number of features increases, however, errors emerge primarily as belief loss: the model struggles to infer the correct posterior from more complicated evidence.
Other models show the opposite pattern in easier settings. They can estimate beliefs reasonably well but fail to apply the specified cost structure when choosing an action. In these cases, optimization residual dominates. Changes in reasoning effort can shift the balance: for some models, additional reasoning reduces failures to follow the loss function, while the main source of error in harder problems remains inaccurate belief formation.
The post-training results reinforce the distinction between the two components. Belief-only training improves belief formation, but that improvement does not necessarily produce better decisions. Decision-only training improves actions without a corresponding improvement in reported beliefs, suggesting that a model can learn a shortcut policy that selects useful actions without forming a more accurate, reusable probability estimate.
Training both components can improve both, but the result depends on how beliefs and decisions are elicited. Joint alignment training is most effective when the training format matches the evaluation format. The authors also report that these findings generalize to loss functions not seen during training, although the strength of the gains depends on the intervention and format.
This focus on the separation between uncertainty reports and downstream behavior connects to work on uncertainty estimates for black-box language models, but the present paper studies a different question: not only whether a model’s probability report is accurate, but whether it converts that belief into an action that respects user-specified costs.

What to keep in mind

The benchmark is deliberately controlled. Its synthetic tasks remove ambiguity about the correct posterior and the intended preferences, which makes diagnosis possible but does not reproduce every complication of real decision support. The weather and clinical experiments add structured text and unstructured notes, yet their ground-truth probabilities still come from specified expert-designed causal models.
The reported regret values are also conditional on the task’s posterior distribution and loss functions. The paper cautions that they should not be interpreted as intrinsic measures of task difficulty or compared indiscriminately across tasks with different underlying distributions.
The central lesson is therefore methodological as much as empirical. A model can make a poor decision because it formed the wrong belief, because it failed to optimize under the stated costs, or because the connection between its belief and action broke down. Improving one stage does not automatically improve the others. For probabilistic decision assistants, evaluating and training these stages separately may be necessary—and ensuring that belief and action are elicited in compatible formats appears particularly important.

Comments (0)

No comments yet

Be the first to share your thoughts!