Back to AI Research

AI Research

Causal Memory Policy tests memories that ordinary retrieval never exposes

Key Takeaways

  • CMP randomizes some retrieval slots to distinguish untested memories from unhelpful ones.
  • Its findings support an exposure audit, while leaving deployment-ready pool selection and
  • Its findings support an exposure audit, while leaving deployment-ready pool selection and prediction of future memory value unresolved.
  • A memory system can mistake “never tested” for “not useful.” If its retriever never puts a stored memory into the language model’s context, adding or deleting that memory may produce the same answer.
  • An estimated effect near zero then says little about what would happen if the model actually read it.

A memory system can mistake “never tested” for “not useful.” If its retriever never puts a stored memory into the language model’s context, adding or deleting that memory may produce the same answer. An estimated effect near zero then says little about what would happen if the model actually read it.
Causal Memory Policy, or CMP, studies that blind spot. Its central intervention changes retrieval exposure, rather than only changing what is stored. The paper also separates estimating a memory’s contribution from deciding whether to retain it.

Why changing the store can leave the experiment unchanged

In the paper’s model, answers depend on the query and retrieved memory texts. Stored memories affect an answer only through retrieval. A memory outside the selected set therefore receives no opportunity to contribute.
This creates a support problem: a comparison needs observations with and without the memory in the model’s context. Randomly changing the store does not supply that comparison if the retriever continues to ignore the memory. The resulting low score can resemble the score of a memory that genuinely does not help.
The distinction matters most for utility-based policies, which rank memories by estimated effects. The authors do not extend the same diagnosis to every retention heuristic. A policy based only on recency or frequency is not making the affected causal-utility estimate.

Reserving context slots for controlled exposure

CMP reserves some context slots for randomized exposure from a specified memory pool. The remaining slots use ordinary ranked retrieval. Across the experiment, exposure slots are distributed as evenly as possible among pool members, with known sampling probabilities.
The method uses self-normalized inverse-propensity weighting to estimate conditional memory utility. Under its balanced design, the stated estimator reduces to a comparison of average outcomes in the two exposure arms. The controlled exposure is what supplies evidence about memories that normal retrieval would miss; changing the estimator alone cannot create missing observations.
Exposure also has a budget. A larger candidate pool needs more timesteps or more reserved slots to give each memory opportunities to appear. The authors note that randomized exposure can select an already retrieved memory, so the realized context may contain fewer distinct memories than the slot budget.

What the results support—and leave open

The paper reports retrieval-level identification failure for 54% of required memories on LongMemEval and 67% on LoCoMo. It also examines a deployed memory system. In its experiments, CMP improves discrimination between required and non-required memories from 0.54 to 0.66 AUC. That is a reported discrimination result, not a percentage of questions answered correctly.
A major qualification concerns the exposure pool. The experiments include memories relevant to the evaluated queries, alongside highly ranked memories and a uniform sample. The authors explicitly say this establishes the effect of randomized exposure, not a deployable selector that reliably finds important memories. Pool selection remains unresolved.
Retention is a separate difficulty. Forgetting removes the opportunity to reconsider a memory later, whereas retaining it can leave that option open. More importantly, the authors report that the tested aggregations of observed-query utility do not predict a memory’s value on unseen queries.
For a memory-system developer, the immediate audit question is whether a low utility score comes from tested exposure or from a retrieval blind spot. CMP offers a controlled way to study that question. It does not establish a complete retention policy that can safely delete memories on the strength of those estimates alone.

Comments (0)

No comments yet

Be the first to share your thoughts!