An agent's earlier experience can help it solve a new task, but it can also carry assumptions that no longer apply. Beyond the Shadows of Plato's Cave proposes FAME, a framework for investigating those memory-related failures through counterfactual reasoning.
The paper uses 'false memory' for inappropriate application of a memory-induced belief, including effects of spurious correlations, environment shifts, and conflicting knowledge. This is a technical evaluation category, not a claim that the agent has human-like recollection or psychological experience.
Asking how a remembered pattern should change
FAME holds the retrieved memory fixed and changes the query to a hypothetical scenario. It then compares the model's internal representation under the original and counterfactual conditions.
A change is not automatically a failure. Some scenarios call for robustness: an irrelevant wording change should not alter the useful concept. Other scenarios require adaptation. If earlier discussions used 'X' as mathematical notation, a later social-media question should elicit a different interpretation.
The framework therefore distinguishes expected stability from expected movement toward a new context. The chosen counterfactual and its admissible region determine how representation drift is interpreted.
Hidden states provide the measurement
The method approximates a memory-induced concept from residual-stream hidden states before answer generation. It uses leave-one-out memory examples to establish reference behavior, then measures the difference produced by a counterfactual query.
The theoretical account includes assumptions about whether demonstrations represent a consistent concept and whether the selected hidden state is sufficient for the predictive distribution. Its answer-readout analysis also uses a specified relationship between hidden states and candidate answer preferences. These conditions matter when transferring the method to another agent architecture.
Avoiding generated answers does not mean avoiding computation. The evaluator still needs model passes, counterfactual construction, and access to the relevant internal representations. A service that exposes only chatbot text would not provide that hidden-state measurement directly.
Detection results are not memory repair
The authors report AUROC values from 76.2% to 96.7% across their false-memory settings. They also report improvements over the best comparison baseline on benchmarks spanning GSM-Symbolic mathematics, GitChameleon code generation, and BigBench-Hard reasoning.
AUROC measures how well a score separates evaluated categories across thresholds. It is not the percentage of all memories that have been corrected, nor a guarantee that every flagged memory causes an incorrect answer. The paper's contribution is evaluation rather than an automatic repair procedure.
The proposed taxonomy gives evaluators a more specific question than whether an answer happened to fail: did the agent apply an old pattern where it needed to adapt, or abandon one where it should have remained stable? The usefulness of that diagnosis depends on the counterfactuals, expected behavior, and model-access assumptions matching the task being studied.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!