EarthVerse is a benchmark designed to evaluate how well scientific AI agents can conduct complex, multi-source investigations into Earth-system events and natural hazards. The researchers behind EarthVerse aim to move beyond simple question-answering tasks by testing whether agents can gather, reconcile, and synthesize heterogeneous data—such as satellite imagery, station records, and impact reports—to form a scientifically reliable and auditable conclusion. Methods and results are detailed in the full paper on arxiv.org.
The Challenge of Scientific Reliability
Current Earth-science benchmarks often provide agents with pre-selected data, testing only their ability to interpret a specific image or passage. EarthVerse changes this by requiring agents to perform "package-scoped investigations." For each of the 405 tasks, an agent is given a collection of files related to a specific disaster—such as a heat wave or flood—and must determine which sources are relevant, reconcile conflicting data, and maintain a consistent chain of evidence throughout their analysis. The researchers note that a single error in scale, unit, or time window can invalidate an entire scientific conclusion, even if individual calculations are correct. The same AI Evaluation question is explored in CAFE, which adds a research perspective.
How the Benchmark Works The benchmark is grounded in 199 documented disasters and extreme events across 19 hazard families. Each task requires the agent to:
Identify evidence: Select relevant files from a package that typically contains about 34 candidates.
Execute calculations: Perform transparent operations like spatial overlaps, ratios, and thresholding.
Maintain provenance: Ensure that every claim is supported by the chosen evidence and that the research process is auditable.
To score performance, the researchers provide "executable ground truth" and task-specific rubrics. They evaluate agents based on both the accuracy of their final answers and the quality of their research process, ensuring that multiple valid paths to a solution are accepted as long as they are scientifically sound. The AI Search story also surfaces in EU Regulators Demand Apple and Google..., adding another angle. The same AI Evaluation question is explored in SciMIF, which adds a research perspective.
Performance and Reliability Gaps In an evaluation of 25 model and agent systems,
the researchers identified a significant "reliability gap." While the best-performing systems achieved a mean answer-unit accuracy of 84.65%, the highest "Strict@95" score—which measures the percentage of tasks completed with near-total accuracy—was only 34.81%. This indicates that while current agents are often capable of completing individual steps of a research task, they frequently struggle to maintain a consistent, error-free chain of reasoning across an entire investigation.
Comments