A retrosynthesis model can recover a recorded reaction while missing a chemist's preference for a different precursor set. Xuemin Chen and colleagues investigate that distinction in Route-Instruction Grounding and Steering, or RIGS. Their method teaches a model which alternatives an instruction favors, then uses that representation to guide precursor generation.
The study concerns single-step retrosynthesis. Although the authors use the word route, it denotes one precursor-set alternative in this paper, rather than a complete multistep synthesis plan. The question is whether a generator can produce several alternatives for the same target product and favor the ones that match a request.
Alternatives have to exist before instructions can select them
Recorded reactions provide limited coverage. The authors construct a multi-route USPTO collection by keeping products with at least two recorded precursor sets; it contains an average of 2.37 observed sets per product. Matching one recorded answer cannot establish that a model understands a preference such as preserving a molecular motif.
The team expands training support with template-derived candidates, using nested budgets of 15, 30, 45 and 60 candidates. Shared graph-construction checks and deduplication keep the comparison consistent. More training alternatives give the generator a wider set of possibilities, but instruction responsiveness needs its own evaluation.
For that test, the authors build guidance with preferred and avoided sets. Their RDKit5 instructions express computable preferences, such as choosing precursor sets with fewer halogen atoms. Labels come from property values calculated within each product's candidate pool. This defines what a successful instruction response means without treating every generated alternative as equally desirable.
Language grounding precedes changes to generation
Stage A trains a language projector against offline comparisons between alternatives for the same product. The language and reaction-pair encoders stay frozen. The projector learns to align an instruction with preferred candidates and separate them from avoided ones.
Stage B connects that representation to a frozen GraphDiT generator through lightweight residual adapters. The model samples precursor sets directly at inference, without test-time retrieval or reranking of a fixed candidate list. Training also limits departures from the unguided model's predictions.
Evaluation distinguishes the ability to generate alternatives from their position in the returned ranking. The authors compare unguided generation, correctly matched instructions and shuffled instructions. That shuffled condition helps test whether an improvement depends on the request's content rather than the mere presence of guidance.
Wider training support does not guarantee stronger control
For the 65M model trained with Top30 support, preferred-set Top-1 recovery rises from 32.93% without guidance to 48.83% with matched instructions. Matched instructions outperform shuffled ones by 18.52 percentage points on that metric. The model also returns fewer avoided candidates in its top ten.
The gains do not keep increasing with the training budget. Top60 weakens instruction-following effects compared with Top30 and Top45, and generating more distinct valid structures does not always recover more candidates in the shared test inventory.
These are controlled model evaluations, not laboratory demonstrations of yield or synthesis feasibility. Molecular structure checks and preference labels measure specific aspects of generation. A chemist would still need to evaluate the proposed chemistry and practical constraints before using an output as an experimental plan.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!