Back to AI Research

AI Research

Can a Cacheable Decision Model Follow Rules? | AI Research

Key Takeaways

  • Can a Cacheable Decision Model Follow Rules?
  • Many AI systems generate an answer and then extract a decision from it.
  • Certo is a small non-generative decision model (Qwen3-4B): it scores candidate actions from their text and returns a probability, instead of generating an answer.
  • The accurate design reads the state, the rules, and each candidate together (a joint scorer), so cost grows with the menu.
  • Independent encoding lets each candidate be encoded once and reused across states (about 5x cheaper at 77 candidates), but separates state from candidate.
Paper AbstractExpand

Certo is a small non-generative decision model (Qwen3-4B): it scores candidate actions from their text and returns a probability, instead of generating an answer. The accurate design reads the state, the rules, and each candidate together (a joint scorer), so cost grows with the menu. Independent encoding lets each candidate be encoded once and reused across states (about 5x cheaper at 77 candidates), but separates state from candidate. We ask how much rule-sensitivity survives that move, and whether it can be trained back. Four experiments on Certo: (1) the tested conversion to cacheable scoring loses rule-sensitivity (recall@1 1.00 -> 0.24) while the joint scorer holds 1.00, and a shortlist+rerank rescue fails; (2) targeted counterfactual supervision restores strong performance on held-out synthetic rule tasks (paraphrase, counterfactual, composition; reproducible across seeds), though we do not isolate whether predictions depend on the supplied rule; (3) on real rules the added benefit is not established -- after fixing a truncation confound, the joint scorer wins significantly on the short tier (0.861 vs 0.500) and directionally on the hard tier (0.655 vs 0.483, n=29); (4) a matched cross-domain real-prose mixture did not help and reduced contract accuracy (-9.3, -16.2 points). A cacheable encoder can be made rule-sensitive on its training distribution, but transfer to unseen-source real rules is not established; the joint scorer keeps an edge at the cost of caching.

Can a Cacheable Decision Model Follow Rules?

Many AI systems generate an answer and then extract a decision from it. Certo takes a narrower approach: it scores a fixed set of candidate actions and returns probabilities, without generating a response. This can make routine decisions faster, but it creates a key design problem. If candidate actions are encoded independently and reused across many situations, can the model still apply the rules that determine which action is actually allowed? The paper studies that tradeoff using a Qwen3-4B-based decision model.

What the two scoring designs do

The reference system is a joint scorer. It reads the situation, the governing rules, and one candidate action together, then assigns that candidate a relevance score. The process is repeated for each option. Because the candidate is interpreted in the context of the state and rules, this design can directly model rule application. Its disadvantage is computational: as the candidate menu grows, more candidate-state combinations must be processed.
The alternative is independent, or cacheable, encoding. The model encodes the state and rules once, encodes every candidate separately, and combines their representations. Candidate vectors can then be reused for other states that share the same menu. In the authors’ measurements, this produced roughly flat latency around 46 milliseconds, while joint scoring rose from 52 milliseconds at five candidates to 242 milliseconds at 77 candidates—about a 5.3-fold difference at the largest tested menu.
This efficiency question is related to systems such as REFLEX’s bounded decision layer for fixed tool choices, although the present paper focuses specifically on whether caching candidate representations harms rule-sensitive selection.

The initial conversion loses rule sensitivity

The first experiment compared several scorers on two tasks: ordinary routing, where an input must be matched to an intent, and rule-sensitive selection, where the correct action depends on an explicit policy. The joint scorer achieved 1.00 recall@1 on the rule-sensitive task at both tested menu sizes.
The cacheable single-vector model performed much worse. Its rule-sensitive recall@1 fell from the joint model’s 1.00 to 0.54 with five candidates and 0.24 with 17 candidates. A cacheable multi-vector model fell further, reaching 0.08 at 17 candidates. Both cacheable systems retained some routing ability, suggesting that they could match broad semantic similarity while struggling to determine which superficially plausible action a rule permitted.
The authors also tested a hybrid rescue: use the cacheable model to retrieve a shortlist, then apply the accurate joint scorer only to that shortlist. This failed because the cacheable stage often removed the compliant candidate before reranking. At shortlist sizes of five and ten, rule-compliant candidate inclusion was only 0.72 and 0.80; it reached 1.00 only when all 17 candidates were retained. A fast first stage cannot recover an answer it has already discarded.

Targeted training helps on synthetic rules

Instead of changing the architecture, the second experiment changed the training data. The single-vector cacheable model received counterfactual examples in which one condition changed and the correct action flipped. These examples deliberately included a semantically tempting but prohibited action as a hard negative.
On held-out synthetic tests, this training produced large gains. Recall@1 reached 0.99–1.00 on paraphrased rules, compared with roughly 0.51–0.55 before the intervention. On counterfactual pairs, the trained model selected both sides correctly 98.7% of the time, versus 30% for the original model. It also handled combinations of rule operators and achieved partial success on rule types absent from training. The findings reproduced under a second random seed.
However, the authors do not treat these results as proof that the model genuinely uses the supplied rule. The synthetic tasks changed facts and conditions in structured ways, so a model might learn a reliable template-specific procedure without comparing the stated rule against the candidates. A decisive control would hold the facts and candidates fixed while changing only the rule, causing the correct answer to change. That rule-only counterfactual test was not run.

Real-world transfer remains unresolved

The third experiment evaluated human-authored insurance, policy, and contract decisions. After correcting a truncation problem, the joint scorer retained a clear advantage on short inputs that fit the models’ context windows: recall@1 was 0.861 for the joint scorer versus 0.500 for the trained cacheable model across 36 items. On a harder fitting subset of 29 items, the joint scorer led 0.655 to 0.483, but the difference was not statistically significant.
The truncation issue mattered because the joint scorer appends each candidate after the document. On long inputs, right-truncation could remove the candidate entirely. Among 38 truncated hard-tier items, every model performed near or below chance, with the joint scorer falling to 0.079. The authors therefore separate fitting and truncated inputs rather than treating the original flat result as evidence about reasoning ability.
Most importantly, the synthetic training did not show a convincing improvement over the original cacheable encoder on these real rules. A further pilot replaced part of the synthetic training mixture with privacy-policy examples, but this reduced accuracy on unseen contract datasets by 9.3 and 16.2 percentage points, depending on the evaluation slice. Because the intervention also removed synthetic examples, the cause of the decline cannot be isolated.
The paper’s conclusion is therefore cautious. Targeted supervision can make a cacheable encoder perform strongly on held-out synthetic rule tasks, while preserving its serving advantage. But the experiments do not establish that it bases those decisions on the supplied rule, nor that the ability transfers to real rules from new sources. For applications where rule sensitivity is critical and menus are small, the joint scorer remains the safer choice. For large menus, caching offers a substantial speed advantage—but with an accuracy cost that must be measured for the specific use case.

Comments (0)

No comments yet

Be the first to share your thoughts!