Frontier AI companies currently use a fragmented set of safety thresholds to determine when a model is too dangerous to release or requires additional security measures. Because these thresholds are often proprietary, incomparable, and lack a shared measurement standard, it is difficult for third parties to audit them or ensure consistent safety across the industry. This paper proposes a methodology to "harmonize" these thresholds by creating a common, auditable floor for safety. By translating different company policies into a shared quantitative framework, the authors aim to prevent a "race to the bottom" where companies might adopt weaker standards to maintain a competitive edge.
A Unified Approach to Risk
The authors categorize AI risks into two distinct types, each requiring a different analytical approach. For misuse risks—specifically cyber and biological threats—the paper argues that "expected harm" should be the primary metric. This involves building explicit risk models that account for how an AI might be used, the conditions under which it is released (such as open weights versus a restricted API), and the potential for the model to increase the success rate of harmful activities. For the third category, automated AI R&D, the authors shift focus away from specific harm pathways and instead propose a method to track the rate of AI progress, identifying when a model’s capabilities break established trends.
Translating Thresholds into Auditable Data
A central contribution of the paper is a translation framework designed to convert the vague or heterogeneous language found in company safety policies into quantitative, auditable floors. In the cyber domain, the authors demonstrate how to map capability-based thresholds (what a model can do) to outcome-based metrics (the expected annual harm). This allows for a more direct comparison between companies. For example, while some companies define thresholds based on "uplift" in attack capabilities, others focus on specific operational outcomes; the proposed framework provides a way to bridge these differences so that independent auditors can verify whether a model has crossed a safety line.
Current Limitations and Future Needs
The authors emphasize that their work is a methodology for comparison and coordination rather than a final set of validated safety numbers. They note that while the framework is robust, the current empirical evidence base is uneven. For instance, while there is enough data to create a functional model for cyber risk, the evidence for biological risk is currently too thin to support a definitive, policy-ready threshold. Furthermore, the authors acknowledge that their initial calibrations rely on author-provided priors. They stress that a mature version of this system would require independent, community-validated data, such as historical incident reports, standardized cyber-range evaluations, and rigorous audits of safety mechanisms.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!