Back to AI Research

AI Research

WorldSolver tests generated physics solvers beyond whether their code runs

Key Takeaways

  • The benchmark combines execution, rendered behavior and numerical physical checks across 168 simulation tasks.
  • A simulation can run without reproducing the physics it was meant to model.
  • [WorldSolver](https://arxiv.org/abs/2610.08720) evaluates coding agents on generating numerical solvers, then checks both their rendered behavior and their underlying state trajectories.
  • The benchmark contains 168 tasks drawn from physical phenomena in 61 computer graphics papers, spanning seven domains.
  • The authors report mean overall scores of 48.7 for GPT-5.6-Sol and 46.7 for Claude-Opus-5 on a 0–100 scale.

A simulation can run without reproducing the physics it was meant to model. WorldSolver evaluates coding agents on generating numerical solvers, then checks both their rendered behavior and their underlying state trajectories.
The benchmark contains 168 tasks drawn from physical phenomena in 61 computer graphics papers, spanning seven domains. The authors report mean overall scores of 48.7 for GPT-5.6-Sol and 46.7 for Claude-Opus-5 on a 0–100 scale. These averages are quality scores, not percentages of completely correct simulations.

Fix the scene and ask the agent to implement its dynamics

Each task provides a description of a physical phenomenon and a code scaffold. Geometry, physical parameters, external inputs and rendering configuration remain fixed. The agent may modify the solver code.
That solver determines how the system's state changes over time. It must export a trajectory containing quantities needed for rendering and physical checks, such as position, velocity or contact forces.
This setup isolates model selection, numerical formulation and implementation from choices about how to decorate or film the scene. Its seven domains include fluids, rigid bodies, elastic solids, cloth and coupled physical systems. A pleasing animation alone does not satisfy the evaluation.

Inspect visible behavior and continuous trajectories

A solver first receives an execution check. If it fails to complete or produces invalid output, its case score is zero. For successful execution, WorldSolver averages visual fidelity and physical plausibility.
A vision-language judge scores sampled video frames for task-specific criteria covering event order, interaction response, the target phenomenon and temporal coherence. The authors use GPT-5.6-Sol as this visual judge and sample 20 key frames per rendered video.
Numerical calculators inspect the state trajectory for relevant laws and constraints. Examples include mass conservation, momentum-impulse consistency and contact nonpenetration. The combination matters because sampled images can miss errors between frames, while a finite set of numerical checks cannot cover every aspect of a phenomenon.
The same need to distinguish executable output from requirement satisfaction appears in A2Z GameSpec-Bench's design-fidelity tests. WorldSolver applies that distinction to physical dynamics.

Read the scores alongside the execution setup

The study evaluates seven frontier models, with Codex CLI for GPT-5.6-Sol, Kimi Code for Kimi-K2.7-Code and Claude Code for the other models. Each agent receives 1,800 seconds for generation, and each generated solver receives a 600-second runtime limit.
The strongest overall models differ by domain. GPT-5.6-Sol leads in fluids, rigid bodies, rods and strands, and multiphysics coupling; Claude-Opus-5 leads in plastic and complex materials, elastic solids, and cloth and shells.
The authors also illustrate a preliminary passing threshold of 60 points, explicitly distinguishing it from an established definition of solver success. The benchmark supports comparing generated solvers under fixed conditions. It does not certify their accuracy for engineering design, and the reported model comparisons also reflect the chosen agent harnesses and visual judging procedure.

Comments