Back to AI Research

AI Research

Cross-model probes test where a language model starts hallucinating

Key Takeaways

  • Hidden-state detectors locate unsupported spans in generated text.
  • The study finds that another model can sometimes detect onsets better than the generator, with important model-ac
  • The study finds that another model can sometimes detect onsets better than the generator, with important model-access and evaluation limits.
  • A language model can begin an answer with accurate information and introduce an unsupported detail halfway through a sentence.
  • Kingshuk Gupta and Davide Buscaldi study whether internal model activations can identify that boundary.

A language model can begin an answer with accurate information and introduce an unsupported detail halfway through a sentence. Kingshuk Gupta and Davide Buscaldi study whether internal model activations can identify that boundary. Their cross-model hallucination-detection paper compares detectors reading a generator's own states with detectors reading the states of another model exposed to its output.
The authors report cases where an external observer matches or exceeds self-detection, including an observer smaller than the generator. That finding concerns locating hallucination onsets in the evaluated data. It does not establish that a second chatbot can verify arbitrary answers without access to model internals.

Locating the start of an unsupported span

The detector assigns each token one of three labels: outside a hallucination, the beginning of a hallucinated span, or continuation of that span. This distinction lets evaluators ask where an error starts rather than assign one score to an entire answer.
The authors first identify attention and feed-forward layers whose average activations differ between factual and hallucinated tokens in calibration data. They then concatenate activations from selected layers into a feature vector for each generated token. These vectors feed supervised classifiers trained against span annotations.
A simple multilayer perceptron evaluates tokens without its own temporal memory. The study also examines sequence models, including bidirectional LSTMs and a BERT-based detector. Conditional random fields constrain the output labels so a continuation cannot appear without a corresponding onset. Those constraints can improve sequence consistency while changing the balance between precision and recall.

Observing another model's text

In the cross-model setting, the researchers extract activation features from the observer's selected layers while it processes the generator's text. The observer provides the representations used for detection; this is not a transfer of a detector between arbitrary hidden-state dimensions.
The experiments include SmolLM2 and TinyLlama generations from PsiloQA, alongside a question-answering subset of RAGTruth using Mistral. The paper compares several observer-generator combinations. Its results support investigating external supervision, but detector architecture and generator choice affect the outcome. More elaborate sequence modeling does not improve onset detection across every generator.
This model-access requirement connects to FAME's hidden-state analysis of agent memory. FAME examines how memory-induced representations change under counterfactual queries; this paper labels unsupported spans in generated text. Both approaches require internal representations, but they diagnose different failure types.

Reading the evaluation under class imbalance

Hallucination beginnings are much rarer than ordinary tokens in the evaluated datasets. A detector can appear successful by favoring the majority class while missing the boundaries that matter.
The paper therefore reports precision-recall diagnostics alongside thresholded precision, recall and F1. Its SmolLM2 comparison finds onset signals above the random prevalence baseline. That result should not be read as a percentage of answers made truthful. The reported comparison also includes an attention-based span detector with comparable intra-model precision-recall performance, so the evidence does not support universal superiority over competing methods.

Deployment questions remain

Some evaluated architectures use bidirectional sequence context. An implementation that must intervene before the next token is delivered would need to establish what information its detector requires and how much delay that adds. The paper's sequence-labeling experiments alone do not settle that operational question.
The work provides a method for measuring error boundaries. Deciding whether to stop generation, retrieve evidence or request human review would require a separate policy and evaluation. For teams building such systems, the useful lesson is to measure onset detection for the particular generator and task instead of assuming the model is its own best observer.

Comments (0)

No comments yet

Be the first to share your thoughts!