Back to AI Research

AI Research

TasteVal measures how efficiently AI researchers choose and interpret experiments

Key Takeaways

  • A new eight-task benchmark separates experimental judgment from coding, comparing models with human experts under the same GPU and time budgets.
  • Choosing the next experiment can matter as much as implementing it.
  • The authors define experimental research taste through compute efficiency: how much serial experimental compute a researcher needs to reach a given score.
  • ## Separating judgment from implementation
  • TasteVal contains eight newly created, open-ended AI research and development tasks.

Choosing the next experiment can matter as much as implementing it. Oliver Jaffe and Dane Sherburn of P-Zero Research examine that decision in TasteVal, a benchmark that asks models to propose experiments, inspect results and decide how to improve a solution. The authors define experimental research taste through compute efficiency: how much serial experimental compute a researcher needs to reach a given score.

Separating judgment from implementation

TasteVal contains eight newly created, open-ended AI research and development tasks. They cover work such as preparing pre-training data, fine-tuning models and modeling human preferences. A model under evaluation acts as the Researcher, while a fixed Coder implements its instructions and reports experimental results. The Researcher cannot inspect the Coder's implementation or the hidden test set.
Each attempt has a forty-H100-hour compute budget and a 120-hour wall-clock limit. GPU-busy time counts toward experimental compute; the Researcher's own token use does not. Experiments run one at a time on a single H100, so the metric measures reaching a score sooner on the same hardware, rather than spreading work over a larger cluster.
The authors recruited twenty-four human experts, with at least two attempting each task. Their expert baseline takes the best human attempt for each task. Humans use the same Coder and experimental budgets, which helps separate differences in research decisions from differences in implementation skill.

Efficiency and final performance tell different stories

The study evaluates twenty models released between 2023 and 2026. The authors report a 2.3x compute multiplier for the strongest model, Opus 5.5, with a 95% confidence interval of 1.15 to 4.37. That multiplier compares the compute needed to reach a matched score relative to the expert baseline. It does not mean every scientific task takes less than half as long.
They estimate that frontier-model compute multipliers doubled about every three months after December 2025, compared with fourteen months before that point. Final normalized performance follows a different pattern: its estimated doubling time is 14.6 months, without a significant trend break. Faster progress toward a score and a higher endpoint are separate measurements.

The boundaries of the result

The tasks favor experiments with quick feedback and verifiable outcomes. The authors withhold task details to reduce contamination, so outside researchers cannot reproduce the entire benchmark from the public paper alone. Opus 5.5 and Fable 5.1 refuse one task; the main analysis treats those refusals as missing observations. Those choices matter when interpreting the headline comparison.
The authors also observe weaker statistical judgment in model transcripts. Human experts measured run-to-run variation early and rejected improvements smaller than that variation, while Fable 5.1 seldom repeated an experiment with a different seed. A successful score therefore leaves room for better uncertainty handling.
TasteVal offers a specific measurement of experimental decision-making under constrained budgets. Its implications for broader research automation depend on whether those tasks resemble the research being automated and whether the measured trends persist. The paper's forecasts remain conditional extrapolations, rather than observed evidence of autonomous scientific discovery.

Comments