Back to AI Research

AI Research

ArchitectureIQ tests whether AI can pick a training recipe before running it

Key Takeaways

  • An execution-grounded benchmark finds strong but uneven training intuition: models handle optimizer comparisons better than architectural choices and often overlook dataset changes
  • An execution-grounded benchmark finds strong but uneven training intuition: models handle optimizer comparisons better than architectural choices and often overlook dataset changes.
  • An AI research assistant can describe why a neural network should train well without correctly predicting which experiment will win.
  • [ArchitectureIQ](https://arxiv.org/abs/2609.39714) measures that pre-execution judgment.
  • A question provides a synthetic dataset and several complete training recipes; the solver must choose the recipe with the best eventual test metric before seeing a run.

An AI research assistant can describe why a neural network should train well without correctly predicting which experiment will win. ArchitectureIQ measures that pre-execution judgment. A question provides a synthetic dataset and several complete training recipes; the solver must choose the recipe with the best eventual test metric before seeing a run.

Ground truth comes from executed experiments

The benchmark contains 500 three-choice questions, divided equally between optimizer-only, architecture-only and mixed comparisons. Each recipe specifies the model, optimizer and hyperparameters, loss, batch size and training duration. The questions include natural-language descriptions and executable Python.
The researchers run the generated experiments to establish the answer. A retained question must have the same winning candidate across ten random seeds. This makes the target more concrete than an expert's preference for one architecture.
Synthetic datasets reduce dependence on familiarity with a named dataset. They also constrain the benchmark's scope: choosing between these short, controlled experiments is different from predicting a large model's behavior on an industrial training corpus.

Strong averages hide weak architectural judgment

The paper reports roughly 76% accuracy for frontier models, above the 33.3% random-choice baseline. Ten human machine-learning researchers each answered a fixed 50-question subset; the best participant reached 66%. The different evaluation sizes matter when interpreting that comparison.
Architecture-only questions expose a weakness. The authors report 38% accuracy for GPT-6 Astra on that category, compared with 65% for the best human participant. Models generally do better when optimizer differences provide a recognizable signal, such as a learning rate too small to move away from initialization.
The researchers also construct paired questions with identical candidate recipes but different datasets. When the dataset change reverses the winning recipe, combined accuracy falls to 46–50%. Models change their selected candidate on only 6–23% of pairs. Those measurements support a specific concern: recipe-level preferences can overwhelm the dataset evidence that should change the decision.

More reasoning versus reusable evidence

Additional inference-time reasoning does not reliably improve results in the paper's fixed 50-question experiment. The authors caution that differences below roughly seven percentage points are not distinguishable at that sample size. Longer explanations alone therefore provide weak evidence of better predictions.
A separate experiment accumulates short propositions from executed outcomes. A solver proposes or invokes a rule, receives evidence from the actual training result, and a curator merges duplicates or narrows overly broad statements. At most twenty propositions enter a subsequent solver prompt.
The paper reports GPT-4o rising from 53.2% without that knowledge base to 67.6% with its final snapshot. That comparison also uses a 50-question subset and does not involve parameter updates. The strongest solver does not show a systematically higher ceiling.
ArchitectureIQ offers a way to test training advice before relying on it. Its useful questions concern which prediction changes with the data, which rule has execution evidence behind it, and whether the proposed experiment still needs to be run. The authors explicitly limit their conclusions to the synthetic tasks, short horizons and finite model families they evaluated.

Comments (0)

No comments yet

Be the first to share your thoughts!