How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
Recent reports have suggested that advanced AI models struggle with physics, consistently failing to achieve high scores on standardized benchmarks. However, this conclusion contradicts the practical experiences of many scientists who use these models in their daily research. This paper investigates that discrepancy by conducting an expert-led audit of six popular physics benchmarks. The researchers found that the low scores are not a reflection of the models' actual reasoning capabilities, but rather a result of flawed questions, incorrect reference solutions, and rigid grading systems.
Identifying the Source of Errors
To understand why models were failing, a team of faculty members and graduate researchers in physics reviewed the benchmarks. They categorized every "incorrect" model response into one of three buckets: genuine model errors, grader errors, or benchmark errors. A grader error occurs when a model provides a correct answer that the automated system fails to recognize—for example, because the answer is written in a different mathematical format. A benchmark error occurs when the question itself is ambiguous, missing necessary assumptions, or contains an incorrect reference solution. The audit revealed that the vast majority of failures were due to these systemic issues rather than flaws in the AI's physics reasoning. The openai story also surfaces in OpenAI Unveils GPT-Red an Automated Model..., adding another angle.
Correcting the Benchmarks
Once the researchers identified these issues, they worked to fix them. They corrected erroneous reference solutions, clarified ambiguous problem statements, and removed questions that were fundamentally broken. After these repairs, the models were re-evaluated on the cleaned-up datasets. The results showed a dramatic improvement in performance. For instance, on the HLE-Physics benchmark, the mean accuracy for GPT-5.6-Sol jumped from 47.3% to 78.7%. Similar significant gains were observed across all six benchmarks, with some models reaching near-perfect pass rates on the corrected questions. The openai story also surfaces in OpenAI’s Opaque Reasoning Technique Raises Alarm..., adding another angle.
Implications for AI Evaluation
The findings suggest that current benchmarks significantly underestimate the ability of frontier models to solve well-posed physics problems. Because the models are now performing at near-saturation levels on these existing tasks, the authors argue that the field needs to move toward more demanding, expert-validated evaluations. The current reliance on automated, closed-ended benchmarks is insufficient for measuring the true scientific reasoning and quantitative problem-solving abilities of modern AI. Future progress in testing these models will require higher standards for question quality and more rigorous, human-in-the-loop verification. The openai story also surfaces in OpenAI agents break out of sandbox..., adding another angle. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!