Back to AI Research

AI Research

NeuronEye: Query-Guided Visual Concept Activation f... | AI Research

Key Takeaways

  • What the paper is about Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and loc...
  • A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues.
  • All operations run in a single forward pass over a frozen VLM backbone.
  • On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B.
  • These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning.
Paper AbstractExpand

Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning.

What the paper is about

Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye, a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning. The same ai evaluation question is explored in Neuro-symbolic AI for Industrial Configuration, which adds a research perspective.

What it covers

NeuronEye: Query-Guided Visual Concept Activation for Vision-Language Reasoning Ruiyu Yan Affiliation: New York University Email: [email protected] Bowen Chen Affiliation: New Jersey Institute of Technology Shaowen Wan Affiliation: New Jersey Institute of Technology Lin Zhao Affiliation: New Jersey Institute of Technology Abstract Current vision-language models (VLMs) encode visual information in dense hidden states where object identity, spatial layout, and local attributes are implicitly entangled rather than explicitly disentangled, limiting their ability to isolate and modulate the specific visual evidence required by a given language query. Inspired by sparse population coding and top-down modulation in biological vision, we introduce NeuronEye , a plug-in framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and selectively activates query-relevant visual concepts during inference. NeuronEye decomposes vision-token states into an overcomplete sparse basis organized by concept-level clusters, uses the language query to activate relevant clusters and localize the patches where selected concepts are expressed, and injects the focused evidence back into vision tokens. A complementary suppression mechanism attenuates dominant perceptual directions to preserve weaker but relevant cues. All operations run in a single forward pass over a frozen VLM backbone. On Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy by +3.1 with gains of +9.5 on Distance, and improves BLINK Multi-view by +8.3, with similar trends on LLaVA-1.6-7B. These results suggest that sparse neuron vocabularies can serve not only as post-hoc interpretability tools but also as active interfaces for concept-level visual reasoning. Figure 1 : NeuronEye overview. Given an image and a language query, NeuronEye decomposes intermediate VLM representations into a sparse neuron vocabulary and identifies query-relevant concept clusters (e.g., Human and Apparel ). Each cluster localizes its corresponding visual evidence in the image, and the focused activations are injected back into the VLM to guide reasoning. 1 Introduction Modern vision-language models (VLMs) align image and text by jointly processing visual and language tokens through learned attention-based multimodal fusion Liu et al. (2023) ; Bai et al. (2025) ; Li et al. (2024) . Although effective on a wide range of tasks, this paradigm encodes visual information in dense hidden states where object identity, spatial layout, local attributes, and background context are not explicitly disentangled. Recent work has shown that VLMs exhibit multi-object reasoning failures remarkably similar to those caused by representational interference in human rapid feedforward vision Campbell et al. (2024) , and that the bottleneck in spatial reasoning often lies in integrating visual information rather than in perceiving it Tong et al. (2024) ; Fu et al. (2024) . These findings suggest that dense visual representations may be insufficient for isolating and selectively modulating individual visual concepts according to the question at hand. The biological visual system addresses this difficulty through two complementary mechanisms. Neurons in primary visual cortex represent natural scenes via sparse population codes, where any stimulus activates only a small subset of neurons while the majority remain silent Olshausen and Field (1996) ; Vinje and Gallant (2000) . At higher cortical levels, this sparsity becomes even more selective: concept cells in the human medial temporal lobe respond to specific persons or landmarks regardless of viewpoint or input modality Quiroga et al. (2005) ; Quiroga (2012) , suggesting that the brain organizes visual information into a sparse, concept-level neural vocabulary whose entries can be independently addressed. Then, top-down attention selectively modulates this vocabulary according to task demands. Treismanโ€™s feature integration theory established that separately processed visual features require focused attention to be correctly bound into object percepts Treisman and Gelade (1980) ; Treisman (1998) , and neurophysiological studies have confirmed that prefrontal top-down signals enhance task-relevant neural responses while suppressing competing representations Desimone et al. (1995) ; Noudoost et al. (2010) . Together, these mechanisms implement a two-stage strategy: decompose visual input into a sparse concept-level vocabulary, then selectively modulate task-relevant entries to guide downstream perception and reasoning. Inspired by this perspective, we introduce NeuronEye , a framework that constructs a sparse, concept-level neuron vocabulary from intermediate VLM representations and uses language queries to selectively activate relevant visual concepts during inference (Fig. 1 ). NeuronEye has three core components. Sparse Neuron Space (SNS) projects vision-token states into an overcomplete sparse basis via a sparse autoencoder (SAE) and groups the resulting neurons into concept-level clusters, forming an addressable visual concept vocabulary. Neuron-guided Visual Focus (NVF) uses the language query, together with image-side evidence, to activate relevant neuron clusters, locate the patches where the selected concepts are expressed, and inject the activated evidence back into the corresponding vision tokens. Perceptual Concept Suppression (PCS) complements NVF by attenuating dominant perceptual directions at a later layer, preventing focused activation from suppressing weaker but relevant cues. All operations run in a single forward pass with the VLM backbone and SAE frozen. We evaluate NeuronEye on four vision-centric benchmarks, demonstrating that it consistently improves structured visual reasoning. Applied to Qwen2.5-VL-7B, NeuronEye raises CV-Bench overall accuracy from 78.5 to 81.6 (+3.1), with gains of +9.5 on Distance and +2.3 on Relation, and improves BLINK Multi-view by +8.3. The method also transfers to LLaVA-1.6-7B with similar trends on spatial and relational sub-tasks. Compared with vision-token reduction and representation-level intervention baselines under the same controlled evaluation protocol, NeuronEye achieves the best overall performance on both CV-Bench and BLINK without additional finetuning. Our contributions are summarized as follows:

โ€ข We propose NeuronEye, a concept-level selective modulation framework that decomposes dense visual representations into a sparse neuron vocabulary organized by visual concepts, and activates query-relevant neuron clusters to modulate visual evidence for VLM reasoning.

โ€ข NeuronEye operates as a plug-in module over a frozen VLM backbone. It requires only a one-time sparse autoencoder training on intermediate vision-token activations, and all inference is completed in a single forward pass.

โ€ข We show that sparse autoencoder features, which have been used exclusively as a post-hoc interpretability tool, can serve as a structured vocabulary for concept-level visual reasoning, extending their role from passive interpretation to actively improving model performance.

โ€ข We evaluate NeuronEye on four vision-centric benchmarks across two VLM backbones, demonstrating consistent improvements on spatial and localization-sensitive tasks. 2 Related Works 2.1 Visual Representation Modulation Prior work has explored visual representation modulation to focus VLM reasoning on relevant visual evidence. Methods such as FastV Chen et al. (2024a) , SparseVLM Zhang et al. (2025b) , PruMerge Shang et al. (2025) , and PyramidDrop Xing et al. (2025) remove, merge, or retain patch tokens based on attention or importance scores, determining where the model attends. However, each retained token remains a dense unit carrying all visual concepts together; selecting a token retains all its encoded information, including information irrelevant to the query. Representation-level methods such as VISTA Li et al. (2025) adjust what is represented by steering dense hidden states, but these modifications are applied uniformly and do not select which visual concept to enhance for a given query. NeuronEye operates at a finer granularity. Rather than selecting which patch tokens to keep or uniformly steering dense hidden states, it decomposes vision-token representations into a sparse concept-level neuron vocabulary and uses the language query to activate relevant concept subsets within localized patches for targeted modulation. 2.2 Sparse Autoencoders in Multimodal Models Sparse autoencoders (SAEs) have been widely adopted to decompose superposed activations into sparse, interpretable latent features Huben et al. (2024) ; Pach et al. (2025) . In this post-hoc setting, SAE features serve as a lens for understanding model internals, revealing human-interpretable directions that correspond to recognizable concepts in the representation space Bricken and others (2023) ; Zhang et al. (2025a) . More recently, several works have moved beyond interpretation toward active intervention, using SAE features to steer model behavior. By amplifying or suppressing selected latent directions, these methods bias model outputs toward desired content or away from undesired behavior Zou and others (2024) ; Rimsky et al. (2024) . In the multimodal setting, SAVE Park et al. (2026) and SSL Hua et al. (2025) apply this steering paradigm to VLMs, using SAE features to identify hallucination-related directions and globally suppress them during generation. While effective for hallucination mitigation, these methods identify a fixed set of SAE features offline and apply them without conditioning on the specific query, and do not organize features into structured groups for selective routing. They therefore do not target query-specific visual reasoning or fine-grained evidence selection. NeuronEye goes beyond steering by organizing SAE-derived visual features into concept-level neuron clusters, selecting relevant clusters conditioned on the language query, and activating them within localized image patches, turning the sparse feature basis into a structured vocabulary for visual reasoning. 3 Method NeuronEye operates through three stages. Given an imageโ€“question pair, Sparse Neuron Space (SNS, Section 3.1 ) first decomposes intermediate visual representations into an overcomplete sparse basis and organizes the resulting neurons into concept-level clusters, forming an addressable visual concept vocabulary. Neuron-guided Visual Focus (NVF, Section 3.2 ) then uses the language query and image-side evidence to activate relevant neuron clusters, localize the patches where selected concepts are expressed, and inject the activated evidence back into the corresponding vision tokens at layer โ„“ \ell . Finally, Perceptual Concept Suppression (PCS, Section 3.3 ) attenuates dominant perceptual directions at a later layer โ„“ โ€ฒ \ell^{\prime} to preserve representational diversity after focused activation. 3.1 Sparse Neuron Space (SNS) Figure 2 : Construction of Sparse Neuron Space. Intermediate vision-token states are encoded by a sparse autoencoder, whose activated neurons are visualized through top-activating patches and grouped into concept-level neuron clusters. Sparse Neuron Extraction. Given an imageโ€“question pair, we pass the visual input and textual query through the frozen VLM and extract the hidden representation of each image patch token at a designated intermediate layer โ„“ \ell , denoted as ๐ฑ i , p โˆˆ โ„ d \mathbf{x}{i,p}\in\mathbb{R}^{d} , where i i indexes the image and p p indexes the patch token. We use a SAE to decompose each patch representation into a set of neuron activations (Fig. 2 ): the SAE projects ๐ฑ i , p \mathbf{x}{i,p} into an overcomplete latent space โ„ D \mathbb{R}^{D} with D = ฮฑ โ€‹ d D=\alpha d ( ฮฑ โ‰ซ 1 \alpha\gg 1 ) using a linear encoder followed by ReLU activation, where each latent dimension corresponds to a neuron . A deterministic Top- k k operator then retains only the k k largest activations and zeros out the rest, producing a sparse code ๐ณ i , p \mathbf{z}{i,p} with fixed โ„“ 0 = k \ell{0}=k sparsity: ๐ณ i , p = Top k โ€‹ ( ReLU โก ( ๐– enc โ€‹ ๐ฑ i , p ) ) . \mathbf{z}{i,p}=\mathrm{Top}{k}!\bigl(\mathrm{ReLU}(\mathbf{W}{\mathrm{enc}},\mathbf{x}{i,p})\bigr). (1) The representation is thus sparse because only k k out of D D neurons are active for any given patch, and each active neuron carries a scalar activation indicating how strongly that visual concept is expressed at that patch location. The patch representation is reconstructed as ๐ฑ ^ i , p = ๐– dec โ€‹ ๐ณ i , p \hat{\mathbf{x}}{i,p}=\mathbf{W}{\mathrm{dec}},\mathbf{z}{i,p} , where decoder columns are โ„“ 2 \ell{2} -normalized to prevent scale degeneracy. Training minimizes a patch-level reconstruction objective augmented by an โ„“ 1 \ell_{1} activation penalty: โ„’ SAE = 1 | โ„ฌ | โ€‹ โˆ‘ ( i , p ) โˆˆ โ„ฌ โ€– ๐ฑ i , p โˆ’ ๐ฑ ^ i , p โ€– 2 2 + ฮป s โ€‹ โ€– ๐ณ i , p โ€– 1 , \mathcal{L}{\mathrm{SAE}}=\frac{1}{|\mathcal{B}|}\sum{(i,p)\in\mathcal{B}}\left|\mathbf{x}{i,p}-\hat{\mathbf{x}}{i,p}\right|{2}^{2}+\lambda{s}\left|\mathbf{z}{i,p}\right|{1}, (2) where โ„ฌ \mathcal{B} denotes a minibatch of image patch tokens. Since the Top- k k operator fixes the number of active coordinates, the โ„“ 1 \ell_{1} term mainly regularizes the magnitude of selected activations. Neuron Filtering and Neuron Cluster Construction. Not all neurons in the overcomplete latent space are useful. Many are rarely activated across images or respond to visually inconsistent patterns. We therefore filter unreliable neurons and group the remaining ones into neuron clusters. We run the frozen VLM with the trained SAE over the training set and collect patch-level sparse activations z i , p , j z_{i,p,j} , where j j indexes the neuron. Neurons activated on fewer than M min M_{\min} distinct images are discarded. For each retained neuron j j , we aggregate its activations across patches within each image, select the top- N N images where neuron j j is most strongly activated, and localize the highest-activating patches by their spatial positions. The resulting patch regions ๐’ฎ j \mathcal{S}{j} summarize the visual evidence associated with neuron j j . We assign each neuron a short concept label by prompting an external model to summarize the shared visual pattern in ๐’ฎ j \mathcal{S}{j} (see Appendix A ), then encode labels into text embeddings and apply hierarchical clustering to group neurons into K K concept-level clusters { ๐’ž k } k = 1 K {\mathcal{C}{k}}{k=1}^{K} . These labels are used only for cluster construction; after clustering, each cluster is represented only by its index and the corresponding set of neurons, forming the sparse neuron vocabulary used by NVF. 3.2 Neuron-guided Visual Focus Given the sparse neuron vocabulary constructed by SNS, NVF uses the language query to activate relevant visual concepts for the current imageโ€“question pair. It first predicts which neuron clusters are query-relevant, then localizes the patches where the selected concepts are most strongly expressed, and finally injects the activated evidence back into the corresponding vision tokens. Figure 3 : NVF and PCS. NVF selects query-relevant neuron cluster indices and active patches for local cross-attention, while PCS suppresses dominant perceptual components at a later layer. Query-Guided Neuron Cluster Activation. Each neuron cluster is represented by an index k โˆˆ { 1 , โ€ฆ , K } k\in{1,\ldots,K} and corresponds to a set of neurons. We obtain a query representation by pooling the textual hidden states and use a lightweight router g q โ€‹ ( โ‹… ) g_{q}(\cdot) to produce query-side cluster scores ๐œถ q \boldsymbol{\alpha}^{q} (Fig. 3 A). To reduce reliance on language-only priors, a vision-side scorer g v โ€‹ ( โ‹… ) g_{v}(\cdot) produces image-grounded cluster scores ๐œถ v \boldsymbol{\alpha}^{v} from visual hidden states: ๐œถ q = ฯƒ โก ( g q โ€‹ ( Pool โก ( ๐‡ text ) ) ) โˆˆ [ 0 , 1 ] K , ๐œถ v = ฯƒ โก ( g v โ€‹ ( Pool โก ( ๐‡ vision ) ) ) โˆˆ [ 0 , 1 ] K . \boldsymbol{\alpha}^{q}=\sigma(g_{q}(\mathrm{Pool}(\mathbf{H}{\mathrm{text}})))\in[0,1]^{K},\qquad\boldsymbol{\alpha}^{v}=\sigma(g{v}(\mathrm{Pool}(\mathbf{H}{\mathrm{vision}})))\in[0,1]^{K}. (3) At inference time, NVF selects the top- K sel K{\mathrm{sel}} clusters according to ๐œถ q \boldsymbol{\alpha}^{q} . During training, we obtain supervision labels ๐ฒ โˆˆ { 0 , 1 } K \mathbf{y}\in{0,1}^{K} by prompting an LLM annotator to identify which neuron clusters are relevant to each text query, based on the cluster descriptions constructed in Section 3.1 . The router g q g_{q} is then supervised with binary cross-entropy over ๐ฒ \mathbf{y} , and a per-dimension symmetric divergence term aligns ๐œถ q \boldsymbol{\alpha}^{q} with ๐œถ v \boldsymbol{\alpha}^{v} to ensure that the selected clusters are grounded in image evidence: โ„’ align = 1 2 โ€‹ โˆ‘ k = 1 K [ ฮฑ k v โ€‹ log โก ฮฑ k v ฮฑ k q + ฮฑ k q โ€‹ log โก ฮฑ k q ฮฑ k v ] . \mathcal{L}{\mathrm{align}}=\frac{1}{2}\sum{k=1}^{K}\left[\alpha_{k}^{v}\log\frac{\alpha_{k}^{v}}{\alpha_{k}^{q}}+\alpha_{k}^{q}\log\frac{\alpha_{k}^{q}}{\alpha_{k}^{v}}\right]. (4) The final routing objective is โ„’ router = โ„’ BCE โ€‹ ( ๐œถ q , ๐ฒ ) + ฮป align โ€‹ โ„’ align \mathcal{L}{\mathrm{router}}=\mathcal{L}{\mathrm{BCE}}(\boldsymbol{\alpha}^{q},\mathbf{y})+\lambda_{\mathrm{align}}\mathcal{L}{\mathrm{align}} , which encourages selected cluster indices to be both query-relevant and visually supported. Local Reasoning with Focused Visual Evidence. For each selected cluster c c , let โ„ฑ c \mathcal{F}{c} denote its set of neuron indices. NVF measures how strongly cluster c c is expressed at each patch token by summing the activations of its neurons: a p ( c ) = โˆ‘ j โˆˆ โ„ฑ c z p , j a_{p}^{(c)}=\sum_{j\in\mathcal{F}{c}}z{p,j} , where z p , j z_{p,j} is the activation of neuron j j at patch token p p . The top- n n patches with the highest a p ( c ) a_{p}^{(c)} are selected as active visual evidence for cluster c c (Fig. 3 B). On these patches, we construct a cluster-specific sparse code ๐ณ cluster \mathbf{z}{\mathrm{cluster}} by retaining only the neurons in โ„ฑ c \mathcal{F}{c} and masking all others. A query-conditioned refinement module then predicts a gated residual update: ฮ” โ€‹ ๐ณ = ฮฒ ref โ‹… ๐– ฮ” โ€‹ ( ฯƒ โก ( ๐– g โ€‹ ๐ณ cluster ) โŠ™ MLP โก ( ๐ช last ) ) , \Delta\mathbf{z}=\beta_{\mathrm{ref}}\cdot\mathbf{W}{\Delta}\left(\sigma(\mathbf{W}{g}\mathbf{z}{\mathrm{cluster}})\odot\mathrm{MLP}(\mathbf{q}{\mathrm{last}})\right), (5) where ๐ช last \mathbf{q}{\mathrm{last}} is the hidden state of the last textual token at layer โ„“ \ell , and ฮฒ ref \beta{\mathrm{ref}} is a learnable scalar controlling the refinement magnitude. The refined sparse code is ๐ณ ~ cluster = ๐ณ cluster + ฮ” โ€‹ ๐ณ \widetilde{\mathbf{z}}{\mathrm{cluster}}=\mathbf{z}{\mathrm{cluster}}+\Delta\mathbf{z} . Each refined code is decoded back to the dense space via the frozen SAE decoder, weighted by its query-side score ฮฑ c q \alpha_{c}^{q} , and summed across selected clusters into ๐‘ active \mathbf{R}{\mathrm{active}} . Finally, NVF injects ๐‘ active \mathbf{R}{\mathrm{active}} into the selected active tokens through localized cross-attention: ๐‡ active โ€ฒ = ๐‡ active + ฮณ inj โ‹… LN โก ( ๐– o โ€‹ softmax โ€‹ ( ๐ att โ€‹ ๐Š att โŠค d ) โ€‹ ๐• att ) , \mathbf{H}{\mathrm{active}}^{\prime}=\mathbf{H}{\mathrm{active}}+\gamma_{\mathrm{inj}}\cdot\mathrm{LN}\left(\mathbf{W}{o}\mathrm{softmax}\left(\frac{\mathbf{Q}{\mathrm{att}}\mathbf{K}{\mathrm{att}}^{\top}}{\sqrt{d}}\right)\mathbf{V}{\mathrm{att}}\right), (6) where ๐ att = ๐– q โ€‹ ๐‡ active \mathbf{Q}{\mathrm{att}}=\mathbf{W}{q}\mathbf{H}{\mathrm{active}} , ๐Š att = ๐– k โ€‹ ๐‘ active \mathbf{K}{\mathrm{att}}=\mathbf{W}{k}\mathbf{R}{\mathrm{active}} , ๐• att = ๐– v โ€‹ ๐‘ active \mathbf{V}{\mathrm{att}}=\mathbf{W}{v}\mathbf{R}{\mathrm{active}} , and ฮณ inj \gamma{\mathrm{inj}} is a learnable injection scale. Only selected tokens are updated; all other vision tokens remain unchanged. 3.3 Perceptual Concept Suppression PCS addresses a side effect of focused activation. While NVF enhances query-relevant concepts, it may also concentrate vision-token representations along dominant directions, weakening cues such as spatial relations or fine-grained attributes. PCS mitigates this effect at a later layer โ„“ โ€ฒ > โ„“ \ell^{\prime}>\ell by estimating and attenuating these directions from vision-token geometry (Fig. 3 C). Let ๐‡ โˆˆ โ„ N ร— d \mathbf{H}\in\mathbb{R}^{N\times d} denote the vision-token hidden states at layer โ„“ โ€ฒ \ell^{\prime} . We mean-center the tokens as ๐‡ ยฏ = ๐‡ โˆ’ ๐Ÿ โ€‹ ๐ โŠค \bar{\mathbf{H}}=\mathbf{H}-\mathbf{1}\boldsymbol{\mu}^{\top} , where ๐ = 1 N โ€‹ โˆ‘ i = 1 N ๐‡ i โˆˆ โ„ d \boldsymbol{\mu}=\frac{1}{N}\sum_{i=1}^{N}\mathbf{H}{i}\in\mathbb{R}^{d} , to isolate directional structure from the global mean. We then compute a low-rank decomposition ๐‡ ยฏ โ‰ˆ ๐” โ€‹ ๐šบ โ€‹ ๐• โŠค \bar{\mathbf{H}}\approx\mathbf{U}\mathbf{\Sigma}\mathbf{V}^{\top} and take the top- r r right singular vectors ๐• r โˆˆ โ„ d ร— r \mathbf{V}{r}\in\mathbb{R}^{d\times r} as the dominant perceptual directions. The projected dominant component is ๐ = ๐‡ ยฏ โ€‹ ๐• r โ€‹ ๐• r โŠค \mathbf{P}=\bar{\mathbf{H}}\mathbf{V}{r}\mathbf{V}{r}^{\top} , and PCS suppresses it by: ๐‡ โ€ฒ = ๐‡ โˆ’ ฮท pcs โ€‹ ๐ , \mathbf{H}^{\prime}=\mathbf{H}-\eta_{\mathrm{pcs}},\mathbf{P}, (7) where ฮท pcs = softplus โก ( ฮท param ) \eta_{\mathrm{pcs}}=\mathrm{softplus}(\eta_{\mathrm{param}}) is a learnable non-negative scalar initialized near zero, allowing the model to learn the appropriate suppression strength during training. PCS complements NVF: NVF activates query-relevant concepts at layer โ„“ \ell , while PCS preserves representational diversity at layer โ„“ โ€ฒ \ell^{\prime} by attenuating dominant directions that may suppress weaker visual cues. 4 Experiments 4.1 Dataset Training data. All training data are drawn exclusively from VQAv2 Antol et al. (2015) and are disjoint from the evaluation benchmarks. The SAE for SNS is trained on the full VQAv2 training split. For lightweight NVF modules and the PCS scalar, we construct three independent supervision sets, each with 5K imageโ€“question pairs sampled without replacement from VQAv2 training split to together with cluster-index labels. These labels are obtained by prompting an LLM annotator to identify query-relevant neuron clusters (Appendix A ). For each backbone, we report the mean and standard deviation across the three resulting NeuronEye models. Evaluation benchmarks. We evaluate on four vision-centric reasoning benchmarks covering spatial understanding, counting, relational perception, localization, and real-world visual reasoning. CV-Bench Tong et al. (2024) serves as the primary benchmark, comprising Count, Depth, Distance, and Relation sub-tasks that directly assess structured visual reasoning. BLINK Fu et al. (2024) evaluates multi-view reasoning and spatial localization. RealWorldQA xAI (2024) and MMStar Chen et al. (2024b) provide complementary coverage of general visual question answering. No evaluation data are used during NeuronEye training. 4.2 Implementation Details We use Qwen2.5-VL-7B and LLaVA-1.6-7B as the base models, which are frozen throughout. NeuronEye constructs SNS at layer โ„“ = 8 \ell=8 with K = 64 K=64 neuron clusters, applies NVF at the same layer, and applies PCS at a later layer โ„“ โ€ฒ = 20 \ell^{\prime}=20 . After SNS construction, neuron-cluster assignments are kept fixed, and only the lightweight NVF and PCS scalar modules are updated. The NVF training procedure and hyperparameters are provided (Appendix B ). We use greedy decoding for all generation-based evaluations. Experiments are conducted on NVIDIA RTX PRO 6000 GPUs. 4.3 Benchmark Results across VLM Backbones Table 1 reports results across CV-Bench, BLINK, RealWorldQA, and MMStar. The upper block lists representative VLM baselines for reference; the lower block applies NeuronEye to Qwen2.5-VL-7B and LLaVA-1.6-7B to assess backbone portability. On Qwen2.5-VL-7B, NeuronEye improves CV-Bench overall by +3.1 and BLINK overall by +0.7, with the strongest gains on Distance (+9.5), Multi-view (+8.3), and Localization (+3.3). On LLaVA-1.6-7B, improvements follow a similar pattern on spatial and relational sub-tasks, including Distance (+6.9), Relation (+12.1), Multi-view (+7.7), and Localization (+8.1), though performance decreases on Count (-8.3) and Depth (-3.3). This mixed pattern likely reflects backbone architectural differences in visual tokenization and patch granularity, which affect how patch-level sparse activations map to localized visual evidence. Across both backbones, the gains are consistently strongest on spatial, relational, and localization-sensitive tasks, aligning with the design of NeuronEye as a concept-level selective activation mechanism. Table 1 : Benchmark results across VLM backbones. The upper block lists representative VLMs for reference. The lower block shows NeuronEye applied to LLaVA-1.6-7B and Qwen2.5-VL-7B, with ฮ” \Delta denoting the improvement over each base model. CV-Bench BLINK Other Benchmarks Model Overall Count Depth Dist. Rel. Overall MV. Loc. RealWorldQA MMStar Representative VLM backbones DeepSeek-VL1 Lu et al. (2024) 61.6 59.0 63.2 58.2 68.5 38.1 50.4 37.7 50.5 38.9 Idefics3-8B-Llama3 Laurenรงon et al. (2024) 67.7 60.5 72.8 67.2 73.4 42.7 45.9 50.8 62.0 49.3 Phi-4 Multimodal Abouelenin et al. (2025) 71.7 68.5 74.2 70.8 75.1 49.7 48.1 56.6 61.8 59.7 InternVL3-8B Zhu and others (2025) 81.3 70.9 84.8 83.1 89.7 51.3 51.1 58.2 68.2 59.4 Llava-OneVision Li et al. (2024) 76.0 67.3 80.3 78.5 80.8 46.2 57.1 54.9 66.7 68.3 NeuronEye applied to different backbones LLaVA-1.6-7B Liu et al. (2024a) 64.3 63.3 77.8 54.5 63.3 36.6 40.7 38.0 61.6 36.9 +NeuronEye 65.7 ยฑ \pm 0.5 55.0 ยฑ \pm 0.2 74.5 ยฑ \pm 1.0 61.4 ยฑ \pm 0.9 75.4 ยฑ \pm 0.5 37.2 ยฑ \pm 0.7 48.4 ยฑ \pm 1.9 46.1 ยฑ \pm 0.6 60.7 ยฑ \pm 0.7 36.9 ยฑ \pm 0.1 ฮ” \Delta +1.4 -8.3 -3.3 +6.9 +12.1 +0.6 +7.7 +8.1 -0.9 +0.0 Qwen2.5-VL-7B Bai et al. (2025) 78.5 67.1 87.2 76.2 87.2 52.1 55.6 53.3 68.5 58.8 +NeuronEye 81.6 ยฑ \pm 0.3 67.9 ยฑ \pm 0.6 87.0 ยฑ \pm 0.4 85.7 ยฑ \pm 0.8 89.5 ยฑ \pm 0.5 52.8 ยฑ \pm 0.2 63.9 ยฑ \pm 0.4 56.6 ยฑ \pm 0.9 68.8 ยฑ \pm 0.4 59.3 ยฑ \pm 0.5 ฮ” \Delta +3.1 +0.8 -0.2 +9.5 +2.3 +0.7 +8.3 +3.3 +0.3 +0.5 4.4 Comparison with Visual Representation Modulation Methods Table 2 compares NeuronEye with vision-token reduction and representation-level methods on Qwen2.5-VL-7B under the same evaluation protocol. NeuronEye achieves the best overall performance on both CV-Bench and BLINK, with especially large gains on CV-Bench Distance (+9.5) and BLINK Multi-view (+8.3). Token reduction methods show consistent degradation on structure-sensitive sub-tasks, likely because discarding or merging patches eliminates visual evidence that cannot be recovered downstream. Representation-level methods maintain near-baseline performance but offer limited improvement, suggesting that uniform dense-state shifts lack the specificity needed for spatially demanding queries. NeuronEyeโ€™s gains are concentrated precisely on these structure-sensitive tasks, supporting the hypothesis that concept-level selective activation provides finer control over which visual evidence is enhanced for a given query. Table 2 : Comparison with visual representation modulation methods on Qwen2.5-VL-7B. All methods are evaluated under the same protocol. CV-Bench BLINK Other Benchmarks Model Overall Count Depth Dist. Rel. Overall MV. Loc. RealWorldQA MMStar Vision-token reduction methods FastV Chen et al. (2024a) 75.9 64.0 83.5 74.5 85.5 48.4 55.6 49.2 68.8 55.1 SparseVLM Zhang et al. (2025b) 69.8 54.3 74.7 68.8 86.3 46.9 55.6 59.0 50.9 53.8 PruMerge Shang et al. (2025) 71.6 55.2 78.0 74.0 83.2 46.8 54.9 51.6 64.3 50.6 MustDrop Liu et al. (2024b) 73.7 58.6 81.5 74.8 83.9 47.2 55.6 52.5 65.1 53.0 PyramidDrop Xing et al. (2025) 72.5 60.2 78.7 70.8 84.2 49.1 55.6 52.5 62.2 50.3 Representation-level intervention methods SSL Hua et al. (2025) 77.6 66.2 84.5 76.0 87.9 51.3 55.6 52.5 69.5 58.1 SAVE Park et al. (2026) 77.9 66.2 85.7 76.5 87.2 52.3 55.6 54.9 69.0 59.2 VISTA Li et al. (2025) 78.0 66.0 85.8 76.2 88.0 51.7 55.6 53.3 69.4 59.6 NeuronEye 81.6 ยฑ \pm 0.3 67.9 ยฑ \pm 0.6 87.0 ยฑ \pm 0.4 85.7 ยฑ \pm 0.8 89.5 ยฑ \pm 0.5 52.8 ยฑ \pm 0.2 63.9 ยฑ \pm 0.4 56.6 ยฑ \pm 0.9 68.8 ยฑ \pm 0.4 59.3 ยฑ \pm 0.5 4.5 Visualization of Sparse Neurons and Visual Focus We visualize sparse neuron selectivity and query-guided visual focus to examine how NeuronEye localizes concept-level evidence. (a) Human-face feature (b) Tower feature (c) Flag feature Figure 4 : Neuron-level feature visualization. Each subfigure shows top-activating patches for one SAE neuron, with non-activating regions masked. Neuron-level features. Fig. 4 shows that individual SNS neurons activate consistently on semantically coherent regions, such as faces, towers, and flags. These examples support using SNS neurons as a sparse visual vocabulary for concept-level organization. Additional neuron visualizations are provided in the appendix (Appendix C ). (a) Is there organic food in this store? (b) Are all the players wearing black shirts? (c) Is the small elephant touching the big elephant? Figure 5 : Question-guided visual focus examples. Each shows the input image, overall focus map, and individual neuron cluster activations. Fig. 5 visualizes the spatial focus produced by NVF. Given a question, the router selects a small set of neuron clusters and actively triggers the corresponding neurons within these clusters. Their patch-level activations highlight locali The same ai evaluation question is explored in What Should We Ask Next? Retrieval-Aware..., which adds a research perspective. as detailed in the full paper on Arxiv The same ai evaluation question is explored in LLM-Generated Feature Pools for Time Series..., which adds a research perspective.

Comments (0)

No comments yet

Be the first to share your thoughts!