Back to AI Research

AI Research

A2Z GameSpec-Bench checks whether generated games follow the designer's complete specifica

Key Takeaways

  • A hundred long-form game-design documents support code, replay, and adaptive-play checks tied to the same requirements.
  • The benchmark distinguishes a game that runs from one that f
  • The benchmark distinguishes a game that runs from one that faithfully implements interacting rules.
  • A generated game can compile and respond to inputs while still violating its designer's rules.
  • [A2Z GameSpec-Bench](https://arxiv.org/abs/2609.39564) evaluates that gap using long-form Game Design Documents, rather than treating a plausible playable result as sufficient.

A generated game can compile and respond to inputs while still violating its designer's rules. A2Z GameSpec-Bench evaluates that gap using long-form Game Design Documents, rather than treating a plausible playable result as sufficient.
The benchmark contains one hundred specifications split between fifty Small and fifty Big designs. Each asks a coding agent to produce a source project and playable build. An agentic authoring pipeline expands briefs into documents and checks them for omissions and inconsistencies.

Requirements form a fixed evaluation contract

Each document becomes a Dependency-Aware Contract containing rules, invariants, and prerequisite relationships. A rule can change a state another rule depends on, so an isolated implementation check may miss a broken gameplay sequence.
The contract is constructed before inspecting the generated build and remains fixed across agents and revision rounds. It is withheld from the building agent during initial generation. This limits the evaluator's ability to quietly redefine success around whatever the agent happened to produce.
The paper illustrates the problem with a card-limit requirement. Source code can appear to implement a removal step, yet playtesting can reveal a reward adding another card before the required removal actually happens. Correct-looking handlers do not establish that their state transitions connect correctly.

Three channels inspect the same design

Source-code inspection checks stated conditions, triggers, effects, and constraints. Scenario-based replay supplies fixed input sequences and collects visual evidence. Adaptive playtesting uses programs that select inputs in response to the running game's state.
Normal play starts from the default state and uses player inputs alone. Targeted adversarial tests can initialize an unverified rule's preconditions, but that assistance is marked in the trace. Subsequent inputs must still produce the specified effects. A prepared starting state is not itself evidence that an outcome occurred.
Judgments and supporting traces remain linked to the same requirements. Rules without an established test situation stay unverified rather than receiving credit simply because no failure was seen.

Running successfully is not design fidelity

The authors report a 98.7% verifiable rate for Claude-Fable-5.1 under compile and runtime checks, alongside an overall GDD Fidelity score of 77.0. These measure different things: the fidelity score is not a percentage of completely correct games.
For the fifty Big designs, requirement-specific feedback improves GDD Fidelity by a reported 10.9% relative to self-revision after two rounds from the same initial builds. That is a relative comparison within this benchmark, not a guarantee for a studio's production workflow.
The main controlled setting uses two-dimensional single-player games in Phaser. The paper also describes an extension to Three.js, but the benchmark should not be read as covering every engine, multiplayer system, or commercial design process.
Its contribution is an evidence structure that connects specifications to implementation and observed behavior. For assessing end-to-end coding agents, passing a build check is a useful starting point; fidelity requires checking the rules together during actual play.

Comments (0)

No comments yet

Be the first to share your thoughts!