Back to AI Research

AI Research

Harmonizing AI Safety Thresholds | AI Research

Key Takeaways

  • Frontier AI companies currently use a fragmented set of safety thresholds to determine when a model is too dangerous to release or requires additional securi...
  • Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies.
  • Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards.
  • We develop a methodology for deriving harmonized thresholds across three risk domains.
  • For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions.
Paper AbstractExpand

Frontier AI companies have published capability thresholds that differ substantially, making it difficult for third parties to verify whether a threshold has been crossed or to compare requirements across companies. Moreover, without common minimum thresholds, risk mitigation may be inconsistent, creating a potential race to the bottom in safety standards. We develop a methodology for deriving harmonized thresholds across three risk domains. For misuse risks (cyber and biological), we take expected harm as the key primitive and use an explicit risk-modeling approach that accounts for risk channels and model release conditions. For automated AI R&D, we base our proposed threshold on the observed rate of AI progress rather than expected harm. Our analysis expands upon prior work and highlights existing empirical gaps and limitations.

Frontier AI companies currently use a fragmented set of safety thresholds to determine when a model is too dangerous to release or requires additional security measures. Because these thresholds are often proprietary, incomparable, and lack a shared measurement standard, it is difficult for third parties to audit them or ensure consistent safety across the industry. This paper proposes a methodology to "harmonize" these thresholds by creating a common, auditable floor for safety. By translating different company policies into a shared quantitative framework, the authors aim to prevent a "race to the bottom" where companies might adopt weaker standards to maintain a competitive edge.

A Unified Approach to Risk

The authors categorize AI risks into two distinct types, each requiring a different analytical approach. For misuse risks—specifically cyber and biological threats—the paper argues that "expected harm" should be the primary metric. This involves building explicit risk models that account for how an AI might be used, the conditions under which it is released (such as open weights versus a restricted API), and the potential for the model to increase the success rate of harmful activities. For the third category, automated AI R&D, the authors shift focus away from specific harm pathways and instead propose a method to track the rate of AI progress, identifying when a model’s capabilities break established trends.

Translating Thresholds into Auditable Data

A central contribution of the paper is a translation framework designed to convert the vague or heterogeneous language found in company safety policies into quantitative, auditable floors. In the cyber domain, the authors demonstrate how to map capability-based thresholds (what a model can do) to outcome-based metrics (the expected annual harm). This allows for a more direct comparison between companies. For example, while some companies define thresholds based on "uplift" in attack capabilities, others focus on specific operational outcomes; the proposed framework provides a way to bridge these differences so that independent auditors can verify whether a model has crossed a safety line.

Current Limitations and Future Needs

The authors emphasize that their work is a methodology for comparison and coordination rather than a final set of validated safety numbers. They note that while the framework is robust, the current empirical evidence base is uneven. For instance, while there is enough data to create a functional model for cyber risk, the evidence for biological risk is currently too thin to support a definitive, policy-ready threshold. Furthermore, the authors acknowledge that their initial calibrations rely on author-provided priors. They stress that a mature version of this system would require independent, community-validated data, such as historical incident reports, standardized cyber-range evaluations, and rigorous audits of safety mechanisms.

Comments (0)

No comments yet

Be the first to share your thoughts!