Finding the right passage is insufficient for questions that require calculating a quantity across records. RECAST trains a language model to choose the operations needed to construct evidence, then gives that evidence to a separate model for the final answer.
The RECAST paper uses a financial example: identifying the three-month period with the largest increase in operating margin requires computing margins and comparing periods. Similarity search might retrieve relevant revenues and operating incomes, while leaving the calculation unfinished.
The router chooses how to obtain evidence
A rule-based preprocessor converts heterogeneous sources into records and creates a compact profile describing their fields and structure. RouterLM receives that profile, the task and the outcomes of earlier operations. It can select lexical retrieval, semantic retrieval or a relational operation, or request customized code.
The primitive operations use BM25, BGE-M3 embeddings and read-only SQLite queries. SQL can filter and join records or calculate aggregates. If those operations cannot express the required transformation, a frozen CompilerLM converts the router's specification into a Python program. Its output returns evidence or an error to the router.
RouterLM decides whether another round is needed. When it accepts the accumulated evidence, a frozen AnswerLM produces the answer in one call. The distinction between successful execution and sufficient evidence remains explicit: code can run without obtaining everything necessary to answer the question.
Training targets the evidence-construction policy
Only RouterLM is trained. The authors first apply supervised fine-tuning to trajectories that produced correct, executable and structurally valid solutions. They then use group relative policy optimization to compare successful and unsuccessful trajectories on the same question.
The reported implementation trains Qwen3.5-9B with LoRA and disables its thinking mode. Gemini 3.5 Flash supplies both the compiler and final answer model. This means the smaller router's result describes a multi-model system, not a standalone 9B model replacing every component.
Across six benchmark families, the trained system reaches mean success of 75.6%, compared with 59.7% for the strongest large-model baseline. Those values differ by 15.9 percentage points. A training-free Gemini router in RECAST reaches 70.6%, making the trained router's advantage five percentage points under this setup.
The comparison has defined budgets and task boundaries
The six families include financial reports, hierarchical tables, Wikipedia passages and user profiles. Evaluation uses 100 unseen questions per family, with three repetitions for in-domain experiments. On three held-out benchmark families, the reported average is 79.3%, compared with 64.3% for the strongest baseline.
All methods share the final answer model and a frozen Gemini judge. Multi-round systems have a six-round limit and retain at most 4,000 characters of evidence. These conditions support controlled comparison, while also limiting what the headline success rate covers. The trained router does not lead every individual benchmark.
For an application with mixed records, RECAST suggests testing the operations that produce the answer's evidence alongside the answer itself. The router's stopping decision, the compiler's implementation and the judge's correctness assessment are separate potential sources of error. Deployment would still require checking those decisions against the actual data and calculations an organization uses.
Comments