Back to AI Research

AI Research

WorldAuditBench tests whether agents can investigate defects in interactive 3D scenes

Key Takeaways

  • WorldAuditBench combines exploration with visual diagnosis across 213 anomaly tasks.
  • Its best evaluated agent reaches 42.3% success, below the human baseline, with especially weak
  • Its best evaluated agent reaches 42.3% success, below the human baseline, with especially weak results on changes that require before-and-after evidence.
  • A solid-looking fence can be defective even when a screenshot looks normal: an agent may be able to walk through it.
  • [WorldAuditBench](https://arxiv.org/abs/2609.40325) evaluates whether multimodal agents can investigate that kind of inconsistency in an interactive 3D environment.

A solid-looking fence can be defective even when a screenshot looks normal: an agent may be able to walk through it. WorldAuditBench evaluates whether multimodal agents can investigate that kind of inconsistency in an interactive 3D environment. The task requires choosing useful actions, gathering evidence and explaining the anomaly, rather than identifying a visual glitch from a supplied image.

Build an audit around observable evidence

The benchmark contains 213 tasks across thirteen environments built with Unreal Engine 5 and Three.js. The authors introduce labeled anomalies into scenes and organize them into five families: static physics, interactive physics, spatial consistency, temporal consistency and semantic consistency. Those families cover fifteen categories.
Some defects are visible in one view, such as an unsupported object. Others require contact, movement or a return visit. Testing a fence's collision behavior requires trying to cross it; identifying a disappearing object requires comparing observations from different times. Each main-evaluation task contains one target anomaly and an evaluation rubric.
Independent judges review the constructed tasks to check that the anomaly can be observed and matches its annotation. A task version needs a Pass rating from at least two judges. That construction process makes this a controlled benchmark with deliberately introduced defects, rather than an estimate of how often commercial games contain bugs.

Compare adaptive exploration with recorded trajectories

The researchers test two auditing arrangements. In the interactive version, a vision-language model chooses actions while inspecting observations and can investigate a suspected problem. In the two-stage version, a vision-language-action model first explores the scene, then another model analyzes its recorded trajectory. The second-stage analyst cannot request new observations once exploration is finished.
The best interactive result in the main table is GPT-6 Astra at 42.3% success. Gemini 3.8 Flash reaches 32.4% and Claude Opus 5 reaches 28.2%; Muse Spark 1.3 and Qwen 3.8 Flash score lower. Two-stage results range from 6.6% to 17.4%. The human baseline reaches 83.4%. Temporal consistency remains difficult: every tested model and arrangement scores at or below 12.5% in that family.
These results support the value of investigating a suspicion while exploring, but the arrangements use different budgets. Interactive agents receive forty environment actions, while the two-stage explorer collects sixty seconds of simulated activity. Human participants have up to ten minutes per task. Those differences belong alongside the headline success rates.

Recognize the limits of the comparison

The study evaluates each model together with its coding-agent harness, using a shared MCP interface for auditing tools. GPT-6 Astra judges reports against the rubrics, including human reports, so the reported scores also depend on that evaluation procedure.
An ablation with Gemini shows that removing both the anomaly-type hint and the matching example reduces interactive success on the Unreal subset from 33.3% to 15.9%. The default task therefore supplies guidance that an unconstrained bug search might lack. WorldAuditBench offers a way to study evidence gathering and diagnosis under defined conditions; its scores do not establish readiness for unrestricted game testing or auditing physical environments.

Comments (0)

No comments yet

Be the first to share your thoughts!