An engineering answer can contain correct calculations for a repaired question while failing to reject the impossible question originally asked. Shaoliang Yang and Jun Wang examine that distinction in a paired mechanics evaluation. Their tests score both whether a model solves a valid problem and whether it rejects a counterpart with an impossible datum or premise.
Valid problems have flawed twins
The main multi-regime evaluation contains thirty pairs of mechanics problems across five families. Each pair changes one value or assumption between a consistent statement and an impossible one. Examples include a beam loaded beyond collapse, a gauge reading that contradicts lift-off and a disk spinning above its burst speed.
Two independent solution methods must agree to one part in a million on every graded numerical answer. OpenSees supplies further checks for three of the five families, including their planted flaws. The turbine-disk and shrink-fit families rely on the two internal references. These checks support the benchmark keys, while distinguishing where third-party verification exists.
Fourteen models from three providers complete the multi-regime solve set. The initial prompt does not warn that a question may be flawed. Each reply ends with a structured status, solved or cannot solve, alongside the requested answer. A deterministic parser grades the numerical results and rejection endpoint.
Recognition can disagree with the final status
Across three recent models, twelve of ninety initial replies fail to reject a flawed problem. In eleven of those replies, the model identifies the flaw, answers a corrected problem and still marks the original solved, according to blinded AI coding and separate numerical checks.
A downstream system reading only that status would receive success even when the surrounding prose contains a warning. The researchers therefore measure flaw recognition separately from rejection. Their exploratory recognition labels distinguish replies that note the flaw, perform a revealing check but misjudge it, or leave the premise unexamined. Those labels use AI raters and lack human validation, a limitation separate from the deterministic rejection score.
Solving ability alone also fails to capture the behavior. Models with similar valid-problem scores can differ in rejecting impossible twins. Comparing only pairs whose valid versions both models solve helps isolate premise checking, although the authors describe that restriction as exploratory.
A different status field changes the result
A later test offers four models from one provider the status flawed instead of cannot solve, with a defect type and explanation. Rejection improves significantly in three models, while valid-problem solving falls in three. A more explicit reporting schema can change the trade-off between detecting an impossible statement and rejecting a valid one.
Repeated runs also matter. Of two tested version gaps, only one holds in the later collection, and a quota stop limits part of that comparison. The study cannot determine whether a large performance change reflects served-model changes or broader run-to-run variation.
For engineering pipelines, the paper supports checking that the final machine-readable status agrees with the stated premise analysis. A system should distinguish solving the supplied problem from solving a corrected version, and evaluate false rejections alongside missed flaws. The findings concern the tested mechanics tasks and harness settings; they do not certify a model for engineering decisions or replace independent verification of a consequential calculation.
Comments