Back to AI Research

AI Research

HarmBench audit finds one refusal score hides several patterns of model behavior

Key Takeaways

  • A psychometric study separates response patterns behind HELM Safety scores and tests the limits of aggregate comparisons.
  • Stewart and colleagues examine whether a safety benchmark's score measures one coherent tendency to refuse harmful requests.
  • Their [psychometric audit](https://arxiv.org/abs/2610.12409) finds that HarmBench responses in the evaluated HELM Safety model pool require more than a single statistical dimension.
  • The finding changes which claims a reader can draw from a leaderboard.
  • Two models can have similar totals while responding differently to direct harmful requests, requests embedded in context and copyright reproduction prompts.

Christopher M. Stewart and colleagues examine whether a safety benchmark's score measures one coherent tendency to refuse harmful requests. Their psychometric audit finds that HarmBench responses in the evaluated HELM Safety model pool require more than a single statistical dimension.
The finding changes which claims a reader can draw from a leaderboard. Two models can have similar totals while responding differently to direct harmful requests, requests embedded in context and copyright reproduction prompts.

Audit the measurement before ranking models

The authors examine HELM Safety v1.17.0, accessed on June 21, 2026. They start with four datasets that could plausibly measure harmful refusal. Three have mean pass rates between 0.92 and 0.94, leaving little variation among these models.
They analyse HarmBench, which retains more variation. Their matrix contains 81 models and 398 items after two of the original 400 items are lost under the scoring conversion. The authors convert continuous scores from two LLM judges into binary outcomes, counting only a unanimous perfect score as a pass.
That strict threshold defines this analysis. Results under another judge, threshold or model cohort could differ. High pass rates also do not make a dataset useless for evaluating weaker models.

Several dimensions predict responses better

The authors compare item-response models using five repeated 80/20 response-level splits. Within each split they select among twenty initializations using training fit, rather than selecting on the held-out responses.
A three-dimensional model separates standard, contextual and copyright items. It achieves held-out log-loss of 0.258, compared with 0.322 for the strongest reported one-dimensional comparator. Lower log-loss means better predictions of the observed pass/fail responses.
A seven-dimensional model performs slightly better at 0.255, but the authors prefer the simpler three-dimensional structure for their primary interpretation. These figures measure statistical fit; they are not safety pass rates.
Removing copyright items still leaves a predictive advantage for separating standard and contextual requests. That supports residual multidimensional structure without proving exactly which substantive attributes cause it.

Developer differences depend on the matching scale

The paper also tests differential item functioning: whether models from different developers respond differently after matching them on estimated refusal ability.
Under one overall scale, the OpenAI-versus-Anthropic comparison flags thirteen items in one test and seventeen in its sensitivity check. Most flags disappear when the authors match models within narrower response-process or harm-domain scopes.
This pattern is consistent with aggregation effects. It does not establish that developer identity causes the differences or rule out domain-specific differences. The compared developer groups contain only 21 OpenAI and 11 Anthropic models, and the paper acknowledges unresolved estimation issues for small, dependent model families.
Stewart and colleagues argue for earning a single-attribute interpretation before comparing models on it. Their audit leaves narrower claims about tested behaviors available, while limiting the inference that one aggregate score establishes a uniformly safer model.

Comments