An Open Pipeline and Dashboard for Systemic-Risk Evidence under the EU AI Act's Code of Practice
The Systemic Risk Index is an open-source evaluation pipeline and interactive dashboard designed to make AI safety evidence transparent and accessible to the public. While the EU AI Act requires providers of general-purpose AI to assess systemic risks, there is currently no standardized way for the public to inspect the evidence behind these claims. This project organizes 19 public benchmarks into four key risk categories—CBRN, cyber offense, harmful manipulation, and loss of control—allowing users to move beyond static, opaque scores and explore the data behind model safety ratings. The same large language models question is explored in LimiX-2, which adds a research perspective.
How the Evaluation Works
The system evaluates 18 different AI models by testing them across three distinct conditions. First, it establishes a baseline using existing benchmarks. Second, it applies "harm-preserving perturbations," such as changing the wording or framing of a query, to see if the model’s safety holds up under different conditions. Finally, it tests the models in simulated agentic scenarios that mimic real-world deployment pressures. By using these varied approaches, the pipeline captures a more nuanced view of model behavior than traditional, single-prompt testing.
Interactive Transparency
The dashboard is built to help non-technical users understand the assumptions that shape AI risk ratings. Users can toggle between "average" and "worst-case" aggregation to see how a model performs when pushed to its limits, rather than relying on a single average score that might hide dangerous failures. Additionally, a capability slider allows users to adjust how much a model’s general intelligence influences its safety score. By clicking on any risk category, users can expand the view to see the specific benchmarks used, read a plain-English description of what is being measured, and trace the rating back to its original source. The same ai safety question is explored in Et Tu, Brute? Economic Misalignment in..., which adds a research perspective.
Key Findings
The research highlights that how data is aggregated significantly changes the perceived safety of a model. When switching from an average assessment to a worst-case aggregation, model scores dropped by 14 to 37 points, demonstrating that average scores often mask significant vulnerabilities. The study also validated its automated grading system, finding that LLM judges agreed with human graders at a level comparable to human-to-human agreement. Furthermore, a user study indicated that participants found the dashboard easy to navigate and felt encouraged to explore how different settings and assumptions impact model safety.
Important Limitations
While the dashboard provides a powerful tool for scrutiny, it is not a definitive certification of safety or legal compliance. The authors note that translating legal risk categories into technical benchmarks involves subjective judgment, and different experts may disagree on which benchmarks best represent specific systemic risks. Additionally, some harms, particularly those related to manipulation, are difficult to test without altering the nature of the request. Because models and their safety safeguards evolve rapidly, the researchers have made their code and data public so that the community can update, rerun, and refine these evaluations as new evidence emerges. The same ai evaluation question is explored in AutoRecLab, which adds a research perspective. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!