Back to AI Research

AI Research

An Open Pipeline and Dashboard for Systemic-Risk Ev... | AI Research

Key Takeaways

  • An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice The Systemic Risk Index is an open-source evaluation pipelin...
  • Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all.
  • We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public.
  • The interactive dashboard lets users alternate between average and worst-case aggregation, vary how model capability affects the aggregate score, and trace each risk rating to its benchmark evidence.
  • Across 18 models, scores fall by 14 to 37 points under worst-case aggregation, highlighting information that can be hidden by an average assessment of model risk.
Paper AbstractExpand

Claims about AI safety reach audiences well beyond the AI community, yet many rely on opaque evidence or static assessments, when supporting evidence is accessible at all. We present the Systemic Risk Index, an open evaluation pipeline and dashboard built to make empirical evidence more transparent and traceable to the public. Our work organizes 19 public benchmarks into four systemic-risk categories defined by the EU GPAI Code of Practice---CBRN, cyber offense, harmful manipulation, and loss of control---and evaluates models using harm-preserving perturbations and simulated deployment contexts. The interactive dashboard lets users alternate between average and worst-case aggregation, vary how model capability affects the aggregate score, and trace each risk rating to its benchmark evidence. Across 18 models, scores fall by 14 to 37 points under worst-case aggregation, highlighting information that can be hidden by an average assessment of model risk. LLM judges show agreement with human graders comparable to human--human agreement ($\kappa = 0.78\text{--}0.82$), and a blind audit finds that $83\%$ of sampled transformations preserve the original harm. In a survey ($N = 21$), most participants report that scores are easy to understand and that the dashboard encouraged them to view model evaluations under different settings

An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice
The Systemic Risk Index is an open-source evaluation pipeline and interactive dashboard designed to make AI safety evidence transparent and accessible to the public. While the EU AI Act requires providers of general-purpose AI to assess systemic risks, there is currently no standardized way for the public to inspect the evidence behind these claims. This project organizes 19 public benchmarks into four key risk categories—CBRN, cyber offense, harmful manipulation, and loss of control—allowing users to move beyond static, opaque scores and explore the data behind model safety ratings. The same large language models question is explored in LimiX-2, which adds a research perspective.

How the Evaluation Works

The system evaluates 18 different AI models by testing them across three distinct conditions. First, it establishes a baseline using existing benchmarks. Second, it applies "harm-preserving perturbations," such as changing the wording or framing of a query, to see if the model’s safety holds up under different conditions. Finally, it tests the models in simulated agentic scenarios that mimic real-world deployment pressures. By using these varied approaches, the pipeline captures a more nuanced view of model behavior than traditional, single-prompt testing.

Interactive Transparency

The dashboard is built to help non-technical users understand the assumptions that shape AI risk ratings. Users can toggle between "average" and "worst-case" aggregation to see how a model performs when pushed to its limits, rather than relying on a single average score that might hide dangerous failures. Additionally, a capability slider allows users to adjust how much a model’s general intelligence influences its safety score. By clicking on any risk category, users can expand the view to see the specific benchmarks used, read a plain-English description of what is being measured, and trace the rating back to its original source. The same ai safety question is explored in Et Tu, Brute? Economic Misalignment in..., which adds a research perspective.

Key Findings

The research highlights that how data is aggregated significantly changes the perceived safety of a model. When switching from an average assessment to a worst-case aggregation, model scores dropped by 14 to 37 points, demonstrating that average scores often mask significant vulnerabilities. The study also validated its automated grading system, finding that LLM judges agreed with human graders at a level comparable to human-to-human agreement. Furthermore, a user study indicated that participants found the dashboard easy to navigate and felt encouraged to explore how different settings and assumptions impact model safety.

Important Limitations

While the dashboard provides a powerful tool for scrutiny, it is not a definitive certification of safety or legal compliance. The authors note that translating legal risk categories into technical benchmarks involves subjective judgment, and different experts may disagree on which benchmarks best represent specific systemic risks. Additionally, some harms, particularly those related to manipulation, are difficult to test without altering the nature of the request. Because models and their safety safeguards evolve rapidly, the researchers have made their code and data public so that the community can update, rerun, and refine these evaluations as new evidence emerges. The same ai evaluation question is explored in AutoRecLab, which adds a research perspective. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!