Jialu Wang and colleagues propose GeoReform, a method for improving the structured geometric information supplied to a multimodal reasoning model. It revises the policy for selecting and presenting diagram relations using feedback from failed solutions.
Their starting observation is that adding geometric facts can both help and harm. On 200 Geometry3K problems, supplying explicit structure to Qwen3VL-8B corrects 28 answers but turns thirteen previously correct answers into errors.
Select relations and ground them to the diagram
GeoReform treats formalization as an intermediate interface. A structure generator receives the diagram, question and current policy, then produces typed relations describing entities, constraints and the target.
The policy controls which information appears and how references map to the diagram. Redundant angle relations can distract a solver from a useful segment equality, while an ambiguous angle label can cause it to apply a constraint to the wrong object.
The method therefore optimizes the organization of facts as well as their extraction. Its structured output omits derivations and extra explanations so that the downstream solver receives a concise representation.
Revise the policy using failed rollouts
During evolution, GeoReform runs the full reasoning pipeline on a training minibatch and gathers failed examples with their diagrams, formalizations, traces and evaluation feedback.
A reflection model proposes a revised policy. The candidate must outperform its parent on the training minibatch before entering the candidate pool and receiving validation evaluation.
The authors use up to twenty update iterations, thirty-two training problems per iteration and a fixed fifty-problem validation minibatch. Final test examples remain held out from evolution.
Only the formalization policy changes. The structure generator, reasoning model and downstream reasoning prompt stay fixed. At test time, the selected policy also stays fixed rather than evolving against the answer to each test problem.
Compare against the correct baseline
The paper reports that Qwen3VL-2B choice accuracy on Geometry3K rises from 42.0% with the initial policy to 56.0% with GeoReform. That is a fourteen-percentage-point gain over the initial policy, rather than the bare model's 33.5% choice accuracy.
Completion accuracy for the same model rises from 37.0% to 53.5%. The distinction matters because completion asks for an answer without candidate options, while choice supplies options.
The experiments also show that explicit structure can lower accuracy in some configurations. On GeoQA, structure helps six of eight reported settings but slightly reduces two. Improving representation quality does not make every added relation useful.
The implementation uses GPT-5 for structure generation and compares downstream Qwen3VL and InternVL3.5 models. These results belong to that pipeline and evaluation budget. They do not establish that any parser or solver will receive the same gains.
GeoReform's useful lesson is specific: a representation can contain true geometric information yet remain difficult for the solver to use. Evaluating the complete reasoning path gives the authors feedback for revising that interface.
Comments