Back to AI Research

AI Research

HARISSA: Inference-Time Self-Checks for Efficient a... | AI Research

Key Takeaways

  • HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment Local language models can preserve privacy, reduce cost, and work...
  • Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models.
  • The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally.
  • A deployment that stays local faces two decisions for hard queries instead.
  • First, it can spend more computation on a query, e.g., reasoning before answering, which raises accuracy at a cost in latency, so it must decide which queries are worth the extra computation (efficiency).
Paper AbstractExpand

Running a language model locally offers advantages in privacy, latency, and cost, but local hardware fits only small models, which are less capable than frontier models. The usual remedy for a hard query, escalating it to a cloud model, gives up the privacy and cost advantages of running locally. A deployment that stays local faces two decisions for hard queries instead. First, it can spend more computation on a query, e.g., reasoning before answering, which raises accuracy at a cost in latency, so it must decide which queries are worth the extra computation (efficiency). Second, some queries are beyond the local model, and delivering a wrong answer is worse than deferring the query to a human in the loop, so it must decide which answers are safe to deliver (safety). We show that both decisions can be made from the model's own hidden states. The prefill state, computed before any token is generated, predicts whether the model will answer correctly, and the answer state, at the end of the generated answer, predicts whether that answer is correct. HARISSA fine-tunes the model so that both states predict correctness, then makes both decisions with one policy that cascades through the ways of answering from cheapest to most expensive, skipping a way the prefill state predicts will fail and deferring the query when the answer it stops with is predicted wrong. On a device running a single model, HARISSA is within one accuracy point of chain-of-thought at 2.7 times lower latency. On a server holding four sizes of one model, HARISSA is more accurate than the FrugalGPT and Self-REF cascades at the same latency, and at the same deferral rate the answer state leaves fewer wrong answers than the standard confidence signals in five of six task and setting pairs.

HARISSA: Inference-Time Self-Checks for Efficient and Safe Local Language Model Deployment

Local language models can preserve privacy, reduce cost, and work without a network connection, but smaller models are more likely to fail on difficult questions. The usual solution—sending those questions to a cloud model—undermines the benefits of local deployment. In this paper on HARISSA, the authors propose using a local model’s own hidden states to decide how much computation a query deserves and whether its final answer is safe to deliver. The system can answer directly, spend more computation on reasoning, try a larger local model, or defer the question to a person.

What HARISSA does

HARISSA addresses two related decisions. The first is efficiency: should the system use a cheap direct answer, or spend more time on chain-of-thought reasoning or a larger model? The second is safety: even after extra computation, is the resulting answer likely to be correct enough to show the user?
The method makes these decisions using two internal signals. The prefill state is the hidden representation at the end of the prompt-processing pass, before the model generates any tokens. HARISSA uses it to estimate whether a particular answering method is likely to succeed. The answer state is the hidden representation at the end of the generated answer. It estimates whether the answer just produced is correct.
This lets the system avoid generating answers that its own prefill state predicts will fail. After an answer is generated, the answer state can cause HARISSA to accept it, discard it and try a more expensive method, or defer the query to a person.
The authors frame these choices with an objective that combines three costs: delivering a wrong answer, spending latency, and deferring a query. In simplified form, the system tries to minimize the expected rate of wrong answers plus weighted latency and weighted deferrals. This makes the trade-off explicit: a deployment may accept some additional computation to reduce errors, but it may also prefer deferral over confidently delivering an unreliable answer.

How the cascade works

HARISSA first fine-tunes the model with LoRA on the deployment’s training data. Every available action answers each training question, and the resulting answers are graded as correct or incorrect. Linear prediction heads are trained on both the prefill and answer states so that they encode these correctness labels.
At inference time, the available actions are ordered from cheapest to most expensive. On a device, the actions are the same model answering directly or using reasoning. On a server, they are four sizes of one model family: 2B, 4B, 9B, and 27B parameters.
For each action except the last, HARISSA first checks the prefill probability. If it falls below a skip threshold, the action is not run at all. Otherwise, the system generates an answer and examines its answer probability. If that probability is high enough, HARISSA stops and keeps the answer. If not, it moves to the next action. At the final action, the system generates an answer and either delivers it or defers it based on the answer probability.
The probes are lightweight logistic regressions over a hidden state, calibrated with isotonic regression. The authors report that each probe requires only one dot product and adds no measurable latency. The approach therefore resembles other confidence-based methods that use uncertainty to control inference, including work on confidence-based control of reasoning efficiency, but HARISSA combines pre-generation routing and post-generation answer checking in one policy.

Results that stand out

The experiments use MedQA, MedMCQA, and sixteen multiple-choice subtasks from BIG-Bench Hard. On a device running Qwen 3.5 4B, HARISSA was within one accuracy point of always using chain-of-thought while taking substantially less time on average. Across the three tasks, the reported average gap was 0.8 percentage points at 2.7 times lower latency.
For example, on MedQA, HARISSA reached 85.7% accuracy at 33.0 seconds per query, compared with 86.2% accuracy at 81.0 seconds for chain-of-thought. On BBH, HARISSA reached 88.7% at 11.7 seconds, while chain-of-thought reached 90.6% at 44.9 seconds. HARISSA also outperformed self-consistency, which samples three direct answers and uses a majority vote, on all three device tasks.
The selective computation was important. On MedQA, HARISSA matched FrugalGPT’s accuracy while using reasoning on 29% of queries, compared with 75% for FrugalGPT. The paper attributes this efficiency to the answer probe identifying incorrect direct answers more effectively than the alternative cascade signals in that setting.
On the server, HARISSA was more accurate than the reported FrugalGPT and Self-REF cascades at the same latency. Its prefill probe can skip model sizes predicted to fail before those models generate an answer, rather than paying the cost of running every smaller model and then escalating.
The answer probability also served as a deferral signal. At the same deferral rate, it left fewer wrong answers delivered than sequence likelihood, logit margin, or Self-REF confidence in five of six task-and-setting comparisons. This is a safety-oriented use of the model’s internal state: the system does not need to answer every question when its own estimate indicates that the answer is unreliable.

What to keep in mind

The reported evaluation is limited to three multiple-choice benchmarks and local deployments built around Qwen models, with device replications on Gemma 4 and Cogito described in the appendix. The device and server settings also assume that a person is available to receive deferred queries. Results may therefore depend on the task, model family, calibration data, and the relative costs assigned to latency and deferral.
HARISSA is not a guarantee that a model can recognize every failure. Its probes are trained from the model’s own correctness patterns, and the final policy depends on thresholds selected on a development split. The paper’s central result is narrower and practical: hidden states can provide useful signals for deciding when to spend more local computation and when not to trust the answer, allowing one deployment to balance speed, accuracy, and deferral without relying on a cloud model.

Comments (0)

No comments yet

Be the first to share your thoughts!