A synthesis planner can propose a plausible-looking chemical reaction and still miss a required reactant or an achievable route to the product. FREA tests how automated feasibility verifiers handle that judgment across several kinds of proposals. Its authors compare model decisions with chemists' labels, rather than assume a recorded reaction is feasible or a generated negative is impossible.
A planning criterion with explicit boundaries
The benchmark contains 751 expert-labeled reactions: 341 feasible and 410 infeasible. Two chemists annotate without seeing the candidate's source or model prediction. They first calibrate their guidelines, label reactions independently and then review disagreements together. The released labels include failure reasons and notes where available.
A feasible proposal must list the substantive reactants required to produce the target. Routine auxiliary components, such as solvents or catalysts, may remain implicit. The criterion also requires reasonable conditions that could yield the target in isolable form through a single step or one-pot operation. It does not require the target to be the major product.
These are expert judgments for planning, not measurements of yield or guarantees of experimental success. That distinction explains an otherwise surprising result: the chemists judge 79 of 93 zero-yield records feasible. Failure under one recorded set of conditions does not exclude a different workable set.
One average hides different weaknesses
FREA draws candidates from retrosynthesis-model proposals, zero-yield records, language-model edits and five methods for generating intended negatives. The study compares prompted language models, forward-prediction models and systems trained on generated negative candidates.
No verifier leads across all four sources. Claude Opus 5 achieves the highest mean balanced accuracy, 68.0, across the sources in this evaluation. Forward models perform best on retrosynthesis proposals, but at the evaluated thresholds they reject many feasible edited reactions. A good screening threshold for one source may discard useful alternatives from another.
The detailed generated-candidate analysis exposes a stronger failure. For alternative product disconnections, Chemformer and BARTSmiles obtain AUROC scores of 26.2 and 32.5. Both are below the chance level of fifty. In that comparison they rank infeasible alternatives above feasible reference reactions, even though they perform better on several other negative-generation methods.
Training gains need a transfer test
The authors also release a corpus containing more than fourteen million recorded reactions and generated negative candidates. Matched training comparisons test whether adding negative supervision helps beyond continued forward-model training.
Negative supervision improves mean AUROC across sources, but the gains do not extend to retrosynthesis proposals. Changing the negative-generation method produces further trade-offs: the largest improvement on generated candidates coincides with worse screening of model proposals. Each training setting uses a single run, so the authors describe those gains as exploratory.
For a chemistry team selecting a verifier, FREA supports evaluating the actual candidate stream and calibrating thresholds on separate representative data. An aggregate benchmark score can help narrow the options, but source-specific errors determine which feasible routes survive and which infeasible proposals enter the next planning step.
Comments