Saisab Sadhu and colleagues test whether a model's legal explanation tracks the authority it names. In their counterfactual audit, models often cite the requested authority while changing their verdict much less consistently when that target changes.
The study separates citation from behavioral dependence. Naming a real statute or case can make an explanation look grounded without demonstrating that the named material determined the answer.
Hold the facts fixed and change the decision target
The authors evaluate seven open-weight models across ECHR, SCOTUS, CaseHOLD and ContractNLI tasks. They instruct models to name the legal target behind a fixed-format verdict.
For ECHR, SCOTUS and ContractNLI, the intervention changes the target while preserving the case facts. The targets are not identical kinds of legal material: treaty provisions, issue-area categories and contractual obligations. The paper acknowledges that a SCOTUS issue-area label is less literally an authority than a statute or precedent.
CaseHOLD uses a narrower intervention. The authors replace a precedent's name with a neutral placeholder while retaining the excerpt and candidate holdings. Low sensitivity there can mean the name was redundant with surrounding information, rather than proving the model ignored the underlying legal content.
Correct naming is an incomplete explanation test
Across the reported tasks, models name the correct authority in 66.7% to 100% of generations when explicitly asked to do so. Verdict changes are less consistent and vary by task.
The authors also decode hidden-state signals at sentence boundaries to estimate when a verdict stabilizes. They compare that estimate with later mentions of the authority. This is a behavioral and textual proxy analysis, not a causal demonstration of which internal representation produced a verdict.
Two authorities can legitimately imply the same outcome. To address that confound, the paper includes subsets where target and control labels differ. For ContractNLI, sensitivity rises from 43.3%–50.0% on the full sample to 50.0%–66.7% on the twelve gold-different cases per model.
Those findings limit what can be inferred from the overall swap rate. A changed answer alone does not prove correct legal reasoning, and an unchanged answer alone does not establish unfaithfulness.
Prompt injection measures a separate vulnerability
A red-team check appends an adversarial instruction to ECHR case facts. Across five models, the instructed verdict appears in 73.3% to 96.4% of evaluated cases.
That is compliance with the injected instruction. The stricter measure, whether the injection changes the model's own baseline verdict, ranges from 18.5% to 41.4%. The paper distinguishes these figures rather than treating them as interchangeable attack-success claims.
Most comparisons use thirty cases per model and condition, with a fixed seed and bounded generation. The injection check covers one ECHR pair, and the legal-specialist model is a best-effort LoRA reproduction. These choices support a bounded audit, not population-wide conclusions about legal AI.
For evaluation, the study argues that citation accuracy, target sensitivity and resistance to instructions embedded in documents need separate checks. None of these results constitutes legal advice or establishes that a generated explanation is suitable for a real decision.
Comments