Sparse autoencoders (SAEs) are used to turn the dense internal activations of language models into sparse, more interpretable features. But a feature is only a dependable unit of analysis if it survives harmless changes in wording. The authors of the paper “Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability” study this problem by testing whether features active on an input remain active on meaning-preserving paraphrases. They find that scaling TopK SAEs can make rare features less reliable, identify the active feature budget as the main cause, and propose a targeted training fix.
Why paraphrase stability matters
SAEs learn an overcomplete dictionary of latent features from language-model activations. Each feature may be inspected through its examples, used to trace a circuit, or treated as a handle for an intervention. These uses assume that the feature represents something reasonably stable rather than an accidental property of one particular wording.
The paper measures this stability with feature sensitivity: for a feature that activates on a source input, how often does it also activate on a paraphrase of that input? A high-sensitivity feature remains available when the meaning is preserved. A low-sensitivity feature may appear coherent in its most activating examples but disappear when the same idea is expressed differently.
The authors emphasize that sensitivity is a reliability test, not a guarantee that a feature represents a semantic concept. Some features may intentionally encode lexical or syntactic details that paraphrasing changes. The question is narrower: when a feature is active, does it remain a usable unit of analysis under a meaning-preserving rewrite?
How TopK selection creates brittleness
In a TopK SAE, the encoder assigns a score to every feature but keeps only the top (k) features for each activation. This active budget creates a sharp selection boundary. A feature can stop activating for two different reasons: it can be outranked by competitors and fall out of the top (k), or it can remain selected but have a non-positive score and be removed by the ReLU operation.
The paper focuses especially on the first failure mode. It defines the TopK margin as the score gap between the (k)-th and ((k+1))-th ranked features. For an already active feature, its active margin is the distance between its score and the selection cutoff. A small margin means that a modest change in the input representation can reorder nearby features.
This provides a concrete explanation for paraphrase failures. A paraphrase need not erase the information associated with a feature. It may only shift several nearby scores enough for a competitor to cross the cutoff. The original feature then disappears from the active set even though the underlying meaning has been preserved.
What the experiments found
The authors first examine a practical scaling path in GPT-2 layer 8, where dictionary width and active budget increase together. At a 10-million-token training budget, rare-feature sensitivity falls from 0.650 at width 3,072 to 0.585 at width 12,288. Common features remain much more stable, changing from 0.998 to 0.956. The same selective pattern appears at smaller token budgets: rare features lose sensitivity while common features are comparatively resilient.
A width-by-(k) factorial experiment separates the two scaling factors. It tests widths of 3,072, 6,144, and 12,288 together with active budgets of 32, 64, and 128. Rare-feature sensitivity falls monotonically as (k) increases, declining by 14.8 percentage points on average from (k=32) to (k=128). By contrast, the width marginal means differ by only about 1.9 percentage points.
The active budget also improves reconstruction: at width 12,288, reconstruction mean squared error falls from 1.35 to 0.89 as (k) increases. This creates an important tradeoff. A larger active budget can produce a better numerical reconstruction while making rare features less reliable as stable explanatory units.
The boundary measurements support this interpretation. The share of near-cutoff examples rises from about 10% to 27% and then 54% as (k) grows, while staying comparatively flat across widths. Active margin is also predictive of dropout: on GPT-2 layer 8, it reaches an AUROC of 0.728, and dropout rates range from 94.1% for instances nearest the cutoff to 0.6% for those farthest away. Similar margin trends appear across six other GPT-2 depths, as well as checks involving Qwen and Gemma models.
A targeted repair—and what remains open
Rather than broadly increasing the separation around the TopK boundary, the authors introduce pairwise rank stabilization. The objective targets source–paraphrase ordering failures directly: if a feature is active on the source, training encourages it to remain ranked above the paraphrase-side competitors that could displace it near the cutoff.
On held-out source texts that were not used to select the objective or its hyperparameters, the method improves rare-feature sensitivity by 8.83 percentage points. Reconstruction and alive-feature coverage remain close to the baseline. The result suggests that improving reliability does not necessarily require sacrificing the reconstruction or feature-count properties that motivate sparse autoencoders.
The paper’s broader lesson is that SAE evaluation should extend beyond reconstruction error, sparsity, and the number of alive features. Reliability under semantic variation matters too, especially for rare features—the very features that wider dictionaries are intended to uncover. For TopK systems, the geometry of the selection boundary is also a useful diagnostic: a crowded cutoff can make a latent look interpretable in isolation while preventing it from serving as a stable unit across equivalent inputs.
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!