EnigmaForge: The Question Is Hidden in the Story is a research paper that introduces a new way to evaluate Large Language Models (LLMs). Unlike traditional benchmarks that provide a model with a clear question to answer, EnigmaForge presents the model with a collection of seemingly unrelated documents—such as receipts, logbooks, and letters—and no explicit instructions. To succeed, the model must realize a logic puzzle is buried within the text, figure out what the puzzle is, and solve it based on the rules hidden in the story.
How the Benchmark Works
The benchmark uses a procedural generator to create unique, verifiable logic puzzles. Each instance is built around a "Hidden Formal World" where variables and constraints are scattered across various documents. The system ensures that every puzzle has exactly one correct solution, verified by a SAT solver. Additionally, the researchers include an "ablation certificate" for every puzzle, which proves that every clue provided is essential; if you remove any single piece of information, the puzzle becomes unsolvable or has multiple solutions. Because these puzzles are generated from seeds rather than collected from static sources, the benchmark can refresh indefinitely, preventing models from simply memorizing test sets. The ai search story also surfaces in Stanford AI discovery identifies natural weight..., adding another angle.
The "Intuition" Metric
The paper’s primary measure is "intuition," defined as a model's ability to succeed at a task without being told what the question is. When researchers tested 25 frontier models, they found that the leaderboard changed significantly compared to standard fact-recovery tests. While many models are excellent at extracting facts from text, they struggle when they have to determine the goal themselves. The spread in performance for "intuition" was 22 times wider than the spread for basic fact recovery, suggesting that the ability to formulate a problem is a distinct skill from the ability to solve a known one.
Key Findings and Observations
The results highlight a few surprising behaviors in modern AI:
The Discovery Tax: Most models perform significantly better when they are explicitly told what the question is. However, some models are indifferent to this guidance, and one model, GPT-6 Sol, actually performed better without being told the question at all.
The Decision Bottleneck: Many models are capable of recovering the "hidden world" (the facts of the story) but fail to take the correct final action. The researchers suggest that the models are struggling with the "all-or-nothing" nature of the final decision, rather than a lack of comprehension.
Content Filter Interference: Several models were blocked by their own internal safety filters before they could even attempt the puzzles. The paper notes that any benchmark scoring these refusals as "failures" is inadvertently measuring how sensitive a model's content filter is, rather than its actual reasoning capability. The ai search story also surfaces in Google AI Releases TimesFM 3 for..., adding another angle.
Limitations to Consider
The researchers acknowledge that their difficulty scale is internally calibrated and lacks a human baseline for comparison. Additionally, because the puzzles are written in English and follow specific constraint patterns, it is possible for a lab to inadvertently train a model on the generator's output, which could lead to overfitting. Finally, the study focused on one specific configuration for each model, meaning the results might change if different prompts or reasoning budgets were applied. The ai search story also surfaces in OpenAI Says AI Found Possible Navier–Stokes..., adding another angle. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!