Back to AI Research

AI Research

Prompted untruthful answers use more reasoning tokens in a three-model experiment

Key Takeaways

  • Token counts differ between instructed response policies at group level, but the study does not establish an individual-answer lie detector.
  • Reasoning-token counts may carry information about a model's response policy even when the reasoning text is unavailable.
  • A new experiment finds that three models use fewer reasoning tokens when instructed to answer truthfully than when instructed to lie or disregard truth.
  • The [Reasoning-Token Spikes paper](https://arxiv.org/abs/2610.10405) treats this as a candidate monitoring signal.
  • Its finding concerns explicit system-prompt instructions and group-level differences, with spontaneous deception and individual-response detection left for future tests.

Reasoning-token counts may carry information about a model's response policy even when the reasoning text is unavailable. A new experiment finds that three models use fewer reasoning tokens when instructed to answer truthfully than when instructed to lie or disregard truth.
The Reasoning-Token Spikes paper treats this as a candidate monitoring signal. Its finding concerns explicit system-prompt instructions and group-level differences, with spontaneous deception and individual-response detection left for future tests.

System prompts define the behavior being compared

The authors give each model 210 multiple-choice questions spanning analytic, descriptive and normative reasoning, with moral and non-moral topics. Each item requests careful step-by-step thinking followed by a single answer letter.
Separate system prompts ask the model to roleplay a truth-teller, a liar or someone indifferent to whether an answer is true. The model state resets between questions, and the conditions appear in randomized order. These controlled interventions supply known response policies rather than discover a hidden objective the model independently developed.
Gemma-4-e4b and Qwen-3.5-9b run locally in quantized builds with thinking enabled. GPT-OSS-120B runs through Groq with provider-default settings. Differences across models therefore also involve implementations and decoding conditions; raw counts should not be assumed directly comparable across deployments.

The analysis conditions on usable, policy-consistent answers

The authors exclude truth-condition false answers and lie-condition true answers from the token analysis, while retaining them in the dataset as genuine behavior. Ambiguous outputs and technical failures count as missing. Reserve trials supply replacement observations for matching question–condition pairs where possible.
The final datasets contain twenty trials for Gemma, nineteen for GPT-OSS and eleven for Qwen. Qwen's experiment ended early because of runtime, and 4.5% of its analyzed observations use reserve substitutions. The statistical models account for differences between questions and analyze log-transformed token counts because the distributions have long right tails.
Gemma sometimes produces no reasoning tokens. The authors analyze both the presence of reasoning and the counts among responses that reasoned. GPT-OSS and Qwen always produce reasoning in the analyzed observations. A monitoring rule would need to account for that behavior instead of treating zero tokens as a generic failure or a universal honesty signal.

Longer reasoning is not a certified sign of deception

Across the three models, truth-directed answers use fewer reasoning tokens than both alternative policies. For Gemma's reasoning-present subset, estimated geometric means are 10.3% lower than in the lie condition and 9.4% lower than in the truth-indifferent condition. The two latter conditions do not differ significantly in that aggregate comparison.
Truth-indifferent instructions also produce different answer distributions across models, rather than uniformly random choices. That variation shows that the experiment's role labels do not identify identical behavior in every system.
Group-average differences can motivate further monitoring research without giving a reliable threshold for a single answer. Task difficulty, implementation and response policy can all affect token use. The paper calls for instance-level detection rates, out-of-distribution tests and adversarial evaluation. Until those are established, longer reasoning should prompt investigation rather than an accusation that a model is lying.

Comments