Back to AI Research

AI Research

BioSecBench-Surveillance: A Verifiable Benchmark fo... | AI Research

Key Takeaways

  • As pathogen genomic surveillance becomes increasingly vital for public health, the primary challenge has shifted from generating data to analyzing it effecti...
  • As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis.
  • We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context.
  • Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically.
  • The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies.
Paper AbstractExpand

As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context. Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically. The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies. Across 3,962 gradable attempts from sixteen model-harness pairs, the strongest configuration cleared only about half. Opus 4.8 with PI led at 50.2 percent, with a 95 percent confidence interval of 40.1 to 60.3 percent across 83 evaluations, tied with GPT-5.5 with Codex at 50.2 percent, with a 95 percent confidence interval of 40.8 to 59.6 percent, followed by Opus 4.7 with PI at 49.6 percent, with a 95 percent confidence interval of 40.0 to 59.2 percent, and Sonnet 4.6 with PI at 48.6 percent, with a 95 percent confidence interval of 38.9 to 58.3 percent. Even when agents invoked the correct workflows, their mistakes came from the choices around them, such as which references, thresholds, filters, and normalization to apply. BioSecBench-Surveillance provides a standard for measuring whether agents can be trusted to perform genomic surveillance when the next outbreak arrives.

As pathogen genomic surveillance becomes increasingly vital for public health, the primary challenge has shifted from generating data to analyzing it effectively. The paper "BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance" introduces a new standardized benchmark designed to evaluate whether AI agents can accurately process raw sequencing data and context to perform genomic surveillance. By providing agents with the same information available to human analysts, the researchers aim to determine if these systems can be trusted to handle the complexities of identifying and tracking pathogens during future outbreaks.

Testing AI in Genomic Analysis

The benchmark consists of 100 distinct evaluations covering seven critical categories, ranging from taxonomic classification to the detection of genetic engineering. These tasks are designed to mirror real-world scenarios, utilizing diverse sample types and various sequencing technologies. To ensure objective measurement, the benchmark uses a deterministic grading system that evaluates the structured answers provided by the AI agents.

Performance of Current AI Models

The researchers tested sixteen different model-harness configurations across 3,962 attempts to see how well they could navigate these surveillance tasks. The results indicate that even the most capable models struggle with the complexity of the work, with the top-performing configurations—Opus 4.8 with PI and GPT-5.5 with Codex—clearing only about 50.2 percent of the evaluations. Other models, such as Opus 4.7 and Sonnet 4.6, followed closely behind, yet none demonstrated a high success rate.

Challenges in Workflow Execution

A key finding from the study is that the difficulty for AI agents often lies in the nuanced decision-making required during analysis. Even when an agent correctly identifies the appropriate workflow to use, it frequently falters in the secondary choices necessary to complete the task. These errors typically involve selecting the wrong references, thresholds, filters, or normalization techniques. This suggests that while AI agents can grasp the high-level requirements of genomic surveillance, they currently lack the precision required for the technical parameters that ensure accurate results.

Comments (0)

No comments yet

Be the first to share your thoughts!