Peter Kulits and colleagues introduce BrickBench, a benchmark asking coding agents to design LEGO assemblies from text prompts. The evaluation separates physical validity, fidelity to the prompt and design quality, because an assembly can pass construction checks without matching a human designer's choices.
The paper supplies BrickAgent, an environment for programmatic construction, inspection and validation. Its results concern agents using those tools, not a language model imagining a build without execution feedback.
Three settings change the construction problem
BrickBench contains 300 prompts, divided among three settings. Model permits at most 400 parts. Set requires between 400 and 4,000 parts. Alt-Build restricts the agent to the inventory of retail set 10698.
The full-library settings use supported LDraw parts. The fixed-inventory setting instead tests whether an agent can produce a design with a scarce, predetermined selection.
BrickAgent includes part search, connector-based placement and subassembly tools. Agents can render designs and receive validation failures identifying the implicated parts, allowing revisions before submission.
A simulated valid build has defined limits
Validity combines part requirements, collision checks and stability tests. The simulator treats connected components as rigid bodies and checks whether they remain at rest under gravity.
That approximation omits effects including stud holding forces and axle play. A design that balances in simulation might still sag or fall apart when assembled. Passing this validator therefore establishes the specified digital checks, not comprehensive physical verification.
The paper reports that leading agents produce valid assemblies for nearly every prompt in the environment. An ablation removing BrickAgent's tools reduces validity for GPT-6 Astra from 100% to 40%. This illustrates the importance of grounded construction feedback within the evaluated setup.
Prompt alignment and design preference are different scores
Semantic alignment uses questions decomposed from the prompt and answered by a vision-language model from rendered views. The best reported overall alignment is approximately 95% of those requirements.
Design quality uses pairwise judgments of rendered assemblies, without supplying the prompt for the design question. The authors convert these comparisons into relative ratings and check agreement with human raters.
The paper reports that human raters identify human-designed assemblies in 323 of 360 comparisons. That supports a remaining design gap within the study, rather than showing that physically valid agent outputs are equivalent to human work.
The human validation covers nine reference agents. Two newer evaluated agents, Claude Opus 5.5 and GPT-6.1 Sol, receive automated scores but were not included in those earlier human studies.
BrickBench demonstrates why a creative benchmark needs more than a pass/fail validator. Its agents can satisfy measurable construction and description constraints while still differing in proportions, part usage and composition. Those design judgments remain dependent on the rendered comparisons and reference set used in the evaluation.
Comments