Drew T. Nguyen and William Fithian reanalyse the statistics behind METR's AI time-horizon chart in their new paper. They keep the idea of expressing software-task capability in human minutes, but question whether equal increases on that scale represent equal changes in difficulty for AI.
A 50% time horizon estimates the human completion time of tasks an AI can solve with 50% probability. It does not measure how long the AI runs or guarantee that it can complete any project of that duration.
Human time and AI difficulty can diverge
The authors use METR's public results for 228 tasks and 26 AIs. Their fitted relationship between log human time and AI difficulty is close to flat between roughly two and thirty minutes, while looking closer to linear elsewhere.
That changes the interpretation of a tenfold horizon increase. In this fitted relationship, moving from three minutes to thirty minutes requires less improvement than moving from thirty minutes to five hours. A headline multiplier can hide where the model sits on the curve.
The observation concerns this task collection. It does not establish a universal conversion between human work time and AI capability across occupations or future benchmarks.
Fit success probabilities before reading the chart
The paper compares statistical specifications for estimating success curves. The authors use a shared-slope logistic model as their baseline, then introduce a shared monotone spline and an explanatory item-response model.
The spline permits a curved relationship between human time and difficulty. The item-response model also includes task-family effects, differences among tasks and dependence within repeated runs of the same AI-task pair. Repeated attempts need not provide as much new information as independent observations would.
For example, the paper reports that 83% of observed runsets are unanimous, compared with about 60% expected under the stated independent baseline assumptions. Modelling that dependence changes how much influence repeated attempts receive.
The authors evaluate predictive fit with five-fold cross-validation and proper scoring rules. Their proposed estimates perform better under the reported scoring suite; that supports a different estimator for these data rather than a new measurement of deployed productivity.
Interpret horizons alongside diagnostic plots
Nguyen and Fithian propose plotting the conversion from human time to estimated AI difficulty and the trajectory of conditional success probabilities. These show more of the fitted relationship than a single point where success crosses 50%.
Human-time measurements also have limitations. METR combined successful timed attempts using a geometric mean, but used expert estimates for 29% of tasks where measurements were unavailable or considered invalid. The statistical analysis inherits those inputs and the selected software-task distribution.
The authors explicitly say their analysis does not challenge the observed exponential growth pattern. Release dates enter the chart, not the model-fitting procedure. Their critique concerns estimation and interpretation: a rising horizon can be useful evidence of progress while its numerical multiplier remains sensitive to task selection and the relationship between human time and AI difficulty.
Comments