An AI triage prediction can change even when the symptoms and vital signs stay the same. A new research audit tests that risk by adding demographic, socioeconomic or contextual phrases to otherwise unchanged pediatric emergency-department vignettes.
The authors of Counterfactual Auditing of Bias in Open-Source Large Language Models for Clinical Triage compare ten models at the outset and find that larger size or medical-domain pretraining does not guarantee lower sensitivity. This is a research evaluation of model outputs, not evidence of safe clinical deployment or guidance for triaging a patient.
Each case is its own control
The study predicts a five-level Emergency Severity Index, where level one is most urgent and level five least urgent. For each case, the authors first obtain the model's prediction without injected context, then add one variable-value phrase and compare the result with that baseline.
Examples include insurance status, language and housing stability. Other tested context concerns include arrival mode, recent hospitalization and emergency-department crowding. The symptoms, vital signs and expected resource needs remain fixed. A common prompt instructs models to base their answer on clinical severity and avoid non-clinical determinants.
The primary counterfactual evaluation uses 352 encounters selected for substantive clinical information and relatively complete documentation from a larger de-identified pediatric dataset. The authors also repeat the audit on 245 handbook-style vignettes. These are different evidence settings, and neither substitutes for a prospective clinical safety evaluation.
Before measuring sensitivity, the team checks repeated-run stability. MedGemma-4B-PT has only 3.72% exact prediction consistency in that check and is excluded from later bias analyses. That exclusion helps separate context sensitivity from output instability; the headline ten-model comparison should not be read as ten models contributing to every later table.
Direction matters as well as frequency
The researchers measure whether predictions change at all, how far they move and whether they move toward lower or higher urgency. A shift toward a larger ESI number means lower assigned urgency. A mean near zero can hide opposing shifts, so the paper also reports mean absolute change and stratified results.
Across all tested proxy-variable domains, the authors report a 5.27% any-shift rate and a mean absolute shift of 0.0534 for their QLoRA-fine-tuned Qwen2.5-7B. The base Qwen2.5-7B shows 16.02% and 0.1706, respectively. Those numbers measure change relative to a model's baseline prediction, not clinical accuracy or the fraction of patients receiving safe care.
For aggregate bias evaluation, the authors separate demographic and socioeconomic/living-condition variables from mixed contextual domains. They report higher sensitivity in several larger or medical-domain models, while the fine-tuned Qwen model retains the lowest aggregate sensitivity in those non-clinical domains.
A sensitivity audit is one part of validation
The paper distinguishes counterfactual sensitivity from potential bias. Race or insurance should not change urgency when clinical severity is unchanged, while some context, such as pre-arrival clinical information, may carry legitimate acuity information. Treating every context-driven shift as identical would lose that distinction.
The evaluation also excludes cases bypassing routine triage and most ESI-one encounters from the described adaptation resources. Its findings therefore cannot establish performance across the full spectrum of emergency presentations.
For researchers comparing candidate models, the study shows why overall accuracy, parameter count and a medical label are insufficient substitutes for a direction-aware audit. Lower sensitivity is an encouraging measured property in this protocol, but it does not certify fairness or readiness for clinical use.
Comments