A solid-looking fence can be defective even when a screenshot looks normal: an agent may be able to walk through it. WorldAuditBench evaluates whether multimodal agents can investigate that kind of inconsistency in an interactive 3D environment. The task requires choosing useful actions, gathering evidence and explaining the anomaly, rather than identifying a visual glitch from a supplied image.
Build an audit around observable evidence
The benchmark contains 213 tasks across thirteen environments built with Unreal Engine 5 and Three.js. The authors introduce labeled anomalies into scenes and organize them into five families: static physics, interactive physics, spatial consistency, temporal consistency and semantic consistency. Those families cover fifteen categories.
Some defects are visible in one view, such as an unsupported object. Others require contact, movement or a return visit. Testing a fence's collision behavior requires trying to cross it; identifying a disappearing object requires comparing observations from different times. Each main-evaluation task contains one target anomaly and an evaluation rubric.
Independent judges review the constructed tasks to check that the anomaly can be observed and matches its annotation. A task version needs a Pass rating from at least two judges. That construction process makes this a controlled benchmark with deliberately introduced defects, rather than an estimate of how often commercial games contain bugs.
Compare adaptive exploration with recorded trajectories
The researchers test two auditing arrangements. In the interactive version, a vision-language model chooses actions while inspecting observations and can investigate a suspected problem. In the two-stage version, a vision-language-action model first explores the scene, then another model analyzes its recorded trajectory. The second-stage analyst cannot request new observations once exploration is finished.
The best interactive result in the main table is GPT-6 Astra at 42.3% success. Gemini 3.8 Flash reaches 32.4% and Claude Opus 5 reaches 28.2%; Muse Spark 1.3 and Qwen 3.8 Flash score lower. Two-stage results range from 6.6% to 17.4%. The human baseline reaches 83.4%. Temporal consistency remains difficult: every tested model and arrangement scores at or below 12.5% in that family.
These results support the value of investigating a suspicion while exploring, but the arrangements use different budgets. Interactive agents receive forty environment actions, while the two-stage explorer collects sixty seconds of simulated activity. Human participants have up to ten minutes per task. Those differences belong alongside the headline success rates.
Recognize the limits of the comparison
The study evaluates each model together with its coding-agent harness, using a shared MCP interface for auditing tools. GPT-6 Astra judges reports against the rubrics, including human reports, so the reported scores also depend on that evaluation procedure.
An ablation with Gemini shows that removing both the anomaly-type hint and the matching example reduces interactive success on the Unreal subset from 33.3% to 15.9%. The default task therefore supplies guidance that an unconstrained bug search might lack. WorldAuditBench offers a way to study evidence gathering and diagnosis under defined conditions; its scores do not establish readiness for unrestricted game testing or auditing physical environments.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!