Which Attention Heads are like the Human Head? Not the Ones that Compute
This paper asks whether attention heads in large language models that resemble human brain activity are also the heads the model needs to solve a task. The authors test that question directly by identifying brain-aligned heads, removing them, and measuring the effect on performance. In an abstract pattern-completion task, they find that brain alignment and computational importance are related only weakly: the most brain-like heads often reflect how a model attends to the input, rather than the mechanisms that are most responsible for producing the answer. Read the paper’s direct test of brain alignment versus causal importance.
What the paper tests
The task involves completing abstract sequences such as AAABAAA → B: the model must infer that the missing item is the distinctive symbol that breaks a run of repeated symbols. The researchers converted the human version of this task into text prompts containing symbols or words, preceded by solved examples. They evaluated LLMs on their immediate next-token answers.
For the human comparison, the authors used fixation-related potentials (FRPs), EEG signals measured around the time a participant looks at an item. They represented both the human signals and each model attention head’s activity as patterns of similarity and dissimilarity across the task’s eight abstract sequence types. A head received a high “brain score” when its representational geometry resembled the frontal EEG geometry.
The study also assigned heads two model-based scores that did not use human data. A patching score estimated how much a head contributed to predicting the correct answer when its activation was transferred from a clean prompt into a corrupted one. Heads ranked highly by this measure were treated as function-vector, or FV, heads. A concept score measured whether a head represented the abstract pattern consistently across different alphabets and response formats, such as English words, German words, symbols, open-ended answers, and multiple-choice answers.
The main analysis began with Llama 3.1 70B and expanded to 16 additional models ranging from 3B to 72B parameters. The cohort included base and instruction-tuned Llama and Qwen models, Phi-4, and DeepSeek-R1-Distill-Llama-70B.
Brain-like did not mean most necessary
The central result is a dissociation between resemblance to human neural activity and causal importance for the model’s answer. In Llama 3.1 70B, brain scores and patching scores were almost unrelated, with a Spearman correlation of 0.03. Across all 17 models, the average correlation was 0.017, and every model showed only a very small relationship.
The authors then removed heads in ranking order and tracked the model’s accuracy. Removing FV-ranked heads was by far the most damaging intervention. At 12.5% of heads removed, FV-ranked ablation caused an average additional accuracy loss of 43 percentage points relative to matched random removals. Concept-ranked removal was also substantially disruptive.
Brain-ranked removal had a smaller effect. It was more damaging than random removal in 15 of the 17 models, showing that brain-aligned heads were not irrelevant. But their causal contribution was much weaker: the largest average excess loss was 12.8 percentage points, and it occurred only after removing about 24.5% of the heads. The result suggests that brain alignment can identify heads carrying some task-related information without identifying the heads on which the computation most strongly depends.
Brain scores were more associated with concept scores than with patching scores, averaging a correlation of 0.10 across models. But this relationship varied considerably. It was consistently positive in the Llama models, sometimes positive in base Qwen models, negative in instruction-tuned Qwen models, and close to zero in Phi-4.
This distinction resembles a broader issue in interpretability: finding units that correlate with a behavior is not the same as showing that they cause it. That concern also appears in work examining whether purported “hallucination neurons” are genuinely causal, although that research studies a different behavior and probing setup.
Two recurring ways of reading the sequence
Among the most brain-aligned heads, the authors found two recurring attention profiles. Novelty heads tended to attend to the distinctive symbol—the item that appears only once in a sequence such as AAABAAA. Repetition heads emphasized positions containing the repeated symbol, such as the As in that same example.
These profiles recurred across the model cohort. A novelty-like cluster appeared in all 17 models when the researchers independently grouped attention patterns, while a repetition-like cluster appeared in 13. The two profiles were strongly opposed: heads that looked more like one template generally looked less like the other.
The profiles differed in their causal effects. Removing repetition heads was more damaging than random removal, with a peak excess loss of 11.2 percentage points. Repetition loading was also positively related to concept scores in Llama 3.1 70B. However, removing repetition heads did not eliminate the model’s measurable representation of the abstract pattern. The authors therefore describe these heads as accompanying or co-varying with concept representation, not necessarily supplying it.
Novelty heads were less important to the tested computations. Removing them was, on average, less damaging than removing an equal number of random heads. This pattern held across seven additional tasks and was also observed when the researchers tracked similar heads during Pythia-1.4B’s training. Their attention profile became more prominent over training, but their measured contribution to the tested computations did not rise to the same degree.
What brain alignment may be capturing
The strongest connection to human behavior appeared in attention rather than answer computation. Human participants changed their gaze allocation depending on the sequence pattern. Across all tested models, novelty-like heads made similar pattern-dependent shifts: they attended to the same distinctive elements that attracted human gaze. Repetition-like heads showed the opposite relationship.
This supports a more limited interpretation of brain alignment. A head can resemble human neural activity because it selects or emphasizes information in a human-like way—especially salient or unusual input—without being one of the model’s essential answer-generating mechanisms.
The authors caution that “novelty” and “repetition” are names for observed attention profiles, not established computational functions. The study also focuses on one abstract sequence-completion setting and a text-based adaptation of the human task. Immediate-answer accuracy does not measure each model’s best performance with extended reasoning, and the relationship between brain alignment and concept representation differs across model families and training variants.
Overall, the findings argue against treating neural similarity as automatic evidence of shared computation. Brain-aligned heads may reveal how a model reads a stimulus—what it notices and where it directs attention—while only faintly revealing how it represents the abstract rule or produces the correct answer.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!