Back to AI Research

AI Research

SciExam for ENSO tests agent-built climate models against hidden scientific checks

Key Takeaways

  • Six of twelve systems exceed a published reference on a composite evaluation, while the reference remains strongest on observed statistics.
  • SciExam for ENSO asks AI agents to build stochastic models of El Niño–Southern Oscillation from real observations.
  • Hidden graders then test their simulations, reconstruction of withheld variables and forecasts, giving the task a scientific evaluation beyond whether an agent completes its code.
  • The [SciExam paper](https://arxiv.org/abs/2610.10513) compares twelve agent systems with a published five-variable model.
  • Six final models exceed the reference's composite score.

SciExam for ENSO asks AI agents to build stochastic models of El Niño–Southern Oscillation from real observations. Hidden graders then test their simulations, reconstruction of withheld variables and forecasts, giving the task a scientific evaluation beyond whether an agent completes its code.
The SciExam paper compares twelve agent systems with a published five-variable model. Six final models exceed the reference's composite score. The advantage comes mainly through reconstruction and forecasting; the reference still reproduces observed ENSO statistics best.

Agents commit to diagnostics before developing the model

Each system works within a shared six-hour budget. It first processes observations into five monthly indices and must pass a data check. It then writes its own diagnostics from scientific guidance about the eventual evaluation. This phase ends when the agent declares completion or reaches 90 minutes.
The diagnostics freeze before model development. Agents can rerun them while revising their models, but cannot change the checks to favor what they have built. Hidden grader scores and held-out data remain unavailable throughout the run. The design lets agents receive feedback without optimizing against the evaluator's disclosed scores.
The models contain central and eastern Pacific sea-surface temperatures alongside thermocline, current and wind variables. Agents choose the deterministic dynamics and stochastic forcing rather than implement a supplied set of governing equations.

Composite scores preserve important differences

The evaluation weights statistical properties and dynamical consistency at 0.3 each, and predictive skill at 0.4. Agents receive observations from 1980–2014; prediction tests use 2015–2024. The reference was developed using observations covering all three evaluation periods and was not tuned to the composite score.
Claude Fable 5.1 leads the reported final scores at 0.515, compared with 0.378 for the reference. These are weighted evaluation scores, not percentages of climate predictions correct. The authors warn against interpreting differences around 0.01 as a reliable ordering among nearby systems.
Runs can also finish below their best checkpoint. Although agents were told their best submission would count, the study compares final checkpoints. Using best checkpoints instead would put eight systems above the reference. That distinction changes the headline and should remain visible.

Candidate mechanisms leave the scientific debate open

The authors simplify nine models approaching or exceeding the reference. Their reduced forms fall into two families compatible with competing explanations of ENSO's warm–cold asymmetry: nonlinear dynamics with additive noise, or linear dynamics with state-dependent noise. Both explanations remain viable under these tests.
The same graders selected the reductions, so improved scores after simplification need caution. One model can reduce into either family with close scores. The benchmark therefore provides candidate structures to inspect, rather than settling the underlying mechanism.
Additional information-access experiments use the top system and run each condition once. Hiding calendar years still yields a competitive model, which weakens a simple recall explanation. Literature access changes its structure, while extra external data can steer calibration away from the assigned observations. These are controlled case studies of one system, not general estimates of how web access affects scientific agents.

Comments