Back to AI Research

AI Research

How Good Are Frontier Models at Physics? Expert Re-... | AI Research

Key Takeaways

  • How Good Are Frontier Models at Physics?
  • Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks Recent reports have suggested...
  • Yet this impression does not always align with domain experts' experiences using these models in their work.
  • We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers.
  • Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning.
Paper AbstractExpand

Low reported scores on leading physics benchmarks, including those featured in the Artificial Analysis Intelligence Index (2026), suggest that frontier language models still struggle with advanced physics, a demanding test of their scientific reasoning and quantitative problem-solving abilities. Yet this impression does not always align with domain experts' experiences using these models in their work. We revisit these reported findings by evaluating frontier models on six widely used physics benchmarks and auditing them with experts, focusing on text-only problems with verifiable final answers. For each subfield of physics, faculty and graduate researchers with relevant expertise carefully review problem statements, reference solutions, and model responses to distinguish genuine model errors from grader errors, incorrect reference solutions, and ambiguous or underspecified questions. Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning. We then ask experts to address these benchmarking issues by correcting erroneous reference solutions and repairing or excluding flawed questions. We find that GPT-5.6-Sol's measured mean@4 rises from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark, while its corrected pass@4 reaches 94.4% on the 54 retained CritPt challenges. Corrected scores are computed on the retained evaluation subsets following expert review. Scores on the audited subsets of UGPhysics, PRISM-Physics, and PHYBench also rise substantially after correction. These findings suggest that current benchmarks substantially understate frontier models' ability to solve well-posed physics problems. Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

Recent reports have suggested that advanced AI models struggle with physics, consistently failing to achieve high scores on standardized benchmarks. However, this conclusion contradicts the practical experiences of many scientists who use these models in their daily research. This paper investigates that discrepancy by conducting an expert-led audit of six popular physics benchmarks. The researchers found that the low scores are not a reflection of the models' actual reasoning capabilities, but rather a result of flawed questions, incorrect reference solutions, and rigid grading systems.

Identifying the Source of Errors

To understand why models were failing, a team of faculty members and graduate researchers in physics reviewed the benchmarks. They categorized every "incorrect" model response into one of three buckets: genuine model errors, grader errors, or benchmark errors. A grader error occurs when a model provides a correct answer that the automated system fails to recognize—for example, because the answer is written in a different mathematical format. A benchmark error occurs when the question itself is ambiguous, missing necessary assumptions, or contains an incorrect reference solution. The audit revealed that the vast majority of failures were due to these systemic issues rather than flaws in the AI's physics reasoning. The openai story also surfaces in OpenAI Unveils GPT-Red an Automated Model..., adding another angle.

Correcting the Benchmarks

Once the researchers identified these issues, they worked to fix them. They corrected erroneous reference solutions, clarified ambiguous problem statements, and removed questions that were fundamentally broken. After these repairs, the models were re-evaluated on the cleaned-up datasets. The results showed a dramatic improvement in performance. For instance, on the HLE-Physics benchmark, the mean accuracy for GPT-5.6-Sol jumped from 47.3% to 78.7%. Similar significant gains were observed across all six benchmarks, with some models reaching near-perfect pass rates on the corrected questions. The openai story also surfaces in OpenAI’s Opaque Reasoning Technique Raises Alarm..., adding another angle.

Implications for AI Evaluation

The findings suggest that current benchmarks significantly underestimate the ability of frontier models to solve well-posed physics problems. Because the models are now performing at near-saturation levels on these existing tasks, the authors argue that the field needs to move toward more demanding, expert-validated evaluations. The current reliance on automated, closed-ended benchmarks is insufficient for measuring the true scientific reasoning and quantitative problem-solving abilities of modern AI. Future progress in testing these models will require higher standards for question quality and more rigorous, human-in-the-loop verification. The openai story also surfaces in OpenAI agents break out of sandbox..., adding another angle. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!