Back to AI Research

AI Research

A neural network that maintains and retrieves memor... | AI Research

Key Takeaways

  • A neural network that maintains and retrieves memories based on context investigates how situational context might shape both working memory and episodic mem...
  • Every day, people continuously infer situational context and adjust the way they understand and remember the world.
  • Context, signaled by the prefrontal cortex, is known to modulate working memory and episodic memory, but the algorithmic understanding of this modulation remains limited.
  • Context also modulates episodic memory retrieval, such that the model retrieves memories based on not only content similarity but also context similarity.
  • This is implemented as a key-value system with self-attention, designed to additionally encode context and retrieve context-congruent memories.
Paper AbstractExpand

Every day, people continuously infer situational context and adjust the way they understand and remember the world. Context, signaled by the prefrontal cortex, is known to modulate working memory and episodic memory, but the algorithmic understanding of this modulation remains limited. Here, we train a recurrent neural network (RNN), augmented with an episodic memory buffer, to infer context using Bayesian inference as it continuously makes predictions of upcoming scenes while watching naturalistic movies. When the inferred context modulates the RNN's recurrent connectivity (the basis of working memory) in a low-rank manner, the model's activity patterns best match neural responses in human participants who watched the same movies during fMRI. Context also modulates episodic memory retrieval, such that the model retrieves memories based on not only content similarity but also context similarity. This is implemented as a key-value system with self-attention, designed to additionally encode context and retrieve context-congruent memories. The resulting model not only better resembles human brain representations but also learns to retrieve memories like humans much faster than a model without context modulation. Together, our findings suggest a computational mechanism by which context modulates information maintenance and long-term memory retrieval in naturalistic environments.

A neural network that maintains and retrieves memories based on context investigates how situational context might shape both working memory and episodic memory. The authors build a recurrent neural network (RNN) that watches television episodes, predicts what will happen next, infers which storyline is currently active, and uses that inferred context to change how it maintains information and retrieves past events. They report that context-sensitive processing makes the model’s internal representations more similar to human brain activity and helps it learn human-like memory retrieval more quickly than a model without context modulation.

What the model does

The model was trained on episodes 2–18 of This Is Us Season 1 and tested on episode 1, which was also watched by 33 human participants during fMRI. The show’s scenes were divided into roughly four-second segments. For each scene, pretrained systems converted video and audio into embedding vectors. The network then tried to predict semantic features of the next scene, such as which characters would appear and where or when the scene would take place.
The model’s “context” was the scene’s storyline. The episode contained four interleaved storylines centered on Jack, Kate, Kevin, and Randall. Some scenes belonged mostly to one storyline, while scenes in which narratives merged could contain more than one context.
During training, the correct context was supplied. At test time, however, the network had to infer it. It simulated predictions under each possible storyline and assigned greater probability to contexts that made better predictions. The previous context estimate served as the next prior, making the inferred context “sticky” across consecutive scenes.

How context changes working memory

The authors first tested where context should influence the RNN. They compared context modulation at the input stage, output stage, and working-memory stage. For working memory, they tested two alternatives: completely separate recurrent connections for each context, and a low-rank mechanism that adjusts a shared recurrent network using a small number of structured changes.
The low-rank model performed best in its comparison with human fMRI representations. The researchers converted both model activity and brain activity into representational similarity matrices, which describe how similarly every pair of scenes was represented. Across the cortex, model–brain correlations were consistently positive. Averaged across cortical regions, the two working-memory conditions outperformed a model with no context modulation, while input and output modulation performed worse than that baseline.
The low-rank working-memory model showed the strongest alignment with the brain. This suggests—within the authors’ modeling framework—that context may influence how information is maintained and processed after it enters the system, rather than simply changing perception at the input or transforming the final output.
This evaluation resembles a broader question in neural-network research: whether brain-aligned internal activity reflects the computations most important for a model’s task. A separate study, Which Attention Heads are like the Human Head? Not the Ones that Compute, examined that relationship directly in language models. The present paper instead uses brain alignment to compare candidate mechanisms for contextual processing.

What the results show

The low-rank model achieved a next-scene prediction correlation of 0.583 ± 0.019 at test, while Bayesian context inference reached 60.61 ± 5.32% accuracy across the four storylines. These results indicate that the network learned enough context-dependent structure to make useful predictions and distinguish the competing contexts, although its inference was far from perfect.
The authors also examined the model’s internal dynamics without visual or audio input. As context was gradually introduced, the low-rank model moved along straighter, lower-dimensional trajectories toward context-specific states. Its context endpoints and recurrent connectivity patterns were also more separated from one another than those of the full-modulation model. In the authors’ interpretation, low-rank gating provides an efficient way to reuse a shared network while still producing distinct states for different contexts.
The model was then extended with an episodic-memory buffer. It stores keys, queries, and values: keys index memories, values contain memory content, and queries search for relevant stored keys. The context-modulated version stores context alongside these representations. During retrieval, it combines content similarity with similarity between the current context and the context in which a memory was stored. The resulting attention weights favor memories that are both content-relevant and context-congruent.
According to the paper, this context-sensitive memory system learned to retrieve memories in a way that matched human participants faster than a model without context modulation—but only when retrieval was selective. This supports the authors’ proposal that context can guide the hippocampus-like memory buffer, rather than merely changing the RNN’s short-term processing.

What to keep in mind

The paper presents a computational mechanism, not a direct demonstration that the brain implements these exact equations. The PFC-like component is represented through Bayesian inference and context-dependent network parameters, rather than as a separately modeled brain region. The authors also note that naturalistic movie paradigms bring confounds and do not provide a perfectly clear objective function.
The study uses storyline identity as its context variable and evaluates the system on one television episode after training on other episodes from the same season. That design creates a realistic sequence of events and long-range dependencies, but it also limits how broadly the findings can be generalized. The central result is therefore best understood as evidence that low-rank context modulation and context-aware episodic retrieval are plausible mechanisms that can jointly reproduce selected behavioral and neural patterns.

Comments (0)

No comments yet

Be the first to share your thoughts!