EarthVerse is a benchmark designed to evaluate how well scientific AI agents can conduct complex, multi-source investigations into Earth-system events and natural hazards. The researchers behind EarthVerse aim to move beyond simple question-answering tasks by testing whether agents can gather, reconcile, and synthesize heterogeneous data—such as satellite imagery, station records, and impact reports—to form a scientifically reliable and auditable conclusion.
The Challenge of Scientific Reliability
Current Earth-science benchmarks often provide agents with pre-selected data, testing only their ability to interpret a specific image or passage. EarthVerse changes this by requiring agents to perform "package-scoped investigations." For each of the 405 tasks, an agent is given a collection of files related to a specific disaster—such as a heat wave or flood—and must determine which sources are relevant, reconcile conflicting data, and maintain a consistent chain of evidence throughout their analysis. The researchers note that a single error in scale, unit, or time window can invalidate an entire scientific conclusion, even if individual calculations are correct.
How the Benchmark Works
The benchmark is grounded in 199 documented disasters and extreme events across 19 hazard families. Each task requires the agent to:
Identify evidence: Select relevant files from a package that typically contains about 34 candidates.
Execute calculations: Perform transparent operations like spatial overlaps, ratios, and thresholding.
Maintain provenance: Ensure that every claim is supported by the chosen evidence and that the research process is auditable.
To score performance, the researchers provide "executable ground truth" and task-specific rubrics. They evaluate agents based on both the accuracy of their final answers and the quality of their research process, ensuring that multiple valid paths to a solution are accepted as long as they are scientifically sound.
Performance and Reliability Gaps
In an evaluation of 25 model and agent systems, the researchers identified a significant "reliability gap." While the best-performing systems achieved a mean answer-unit accuracy of 84.65%, the highest "Strict@95" score—which measures the percentage of tasks completed with near-total accuracy—was only 34.81%. This indicates that while current agents are often capable of completing individual steps of a research task, they frequently struggle to maintain a consistent, error-free chain of reasoning across an entire investigation.
Why This Matters
Earth-system analysis is essential for disaster response and climate risk assessment, where accurate estimates of severity and exposure are critical. The researchers argue that for AI to be useful in these fields, it must be capable of more than just isolated reasoning; it must be able to construct a coherent evidence base from messy, real-world data. EarthVerse provides a reproducible framework to measure this end-to-end reliability, helping to identify specific failure points in how agents access evidence, select tools, and manage long-horizon scientific workflows.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!