Back to AI Research

AI Research

KV-Kaizen: Learning Context-Adaptive Cache Compress... | AI Research

Key Takeaways

  • KV-Kaizen: Learning Context-Adaptive Cache Compression Choices addresses a central deployment problem for long-context large language models: the KV cache ca...
  • As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights.
  • This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size.
  • Recent work alleviates this bottleneck by discarding the least relevant tokens.
  • Eviction introduces a tension, since a one-off decision to discard content may prove detrimental later.
Paper AbstractExpand

As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work alleviates this bottleneck by discarding the least relevant tokens. Eviction introduces a tension, since a one-off decision to discard content may prove detrimental later. Instead, we focus on alternative choices that can lead to cache compression without evicting tokens. We achieve this by learning a selector that is able to produce, based on context, a per-layer cache configuration towards an overall compression budget. The selector operates along three axes: sharing one cache across layers (depth), caching at fewer bits (precision), or truncating the low-rank latent cache representations (rank). We call the resulting method KV-Kaizen, for the many small per-layer choices it compounds. We observe that these interventions taken independently and uniformly over all layers limit achievable compression because they degrade accuracy. Crucially, composing them locally and adaptively to the context can instead preserve accuracy while achieving large memory savings. At inference, the selector runs once, before pre-fill. In evaluations on instruction following and reasoning tasks, our selectors reach the Pareto frontier of accuracy against cache size, against learning-free and post-hoc baselines. On long-context tasks, KV-Kaizen improves on eviction and can be composed with it, reaching a 32x smaller decode-time cache on a 14B model while preserving accuracy. A 4x cache size reduction incurs no accuracy degradation from 7B parameters up, and a compressed model is more accurate than a smaller uncompressed one with the same cache size. Together, these findings support pre-training large models and compressing them only afterwards.

KV-Kaizen: Learning Context-Adaptive Cache Compression Choices addresses a central deployment problem for long-context large language models: the KV cache can become larger than the model weights themselves. Because every generated token must read that cache, decoding is often limited by memory bandwidth. The paper proposes learning how to compress the cache differently for each layer and each input context, without removing tokens. Its reported results show that several mild compression choices can be combined to reduce memory substantially while preserving accuracy.

What the paper does

During generation, an LLM stores key and value vectors for every token in a KV cache. This avoids recomputing earlier tokens, but the cache grows with context length and must be read repeatedly during decoding. For example, the authors note that serving 32 sequences of 8,000 tokens with Qwen2.5-14B can require 48 GiB of cache memory in bf16.
A common response is token eviction: retain positions judged important and discard the rest. The authors argue that this creates a risk because a token that appears unimportant for one future query may later become useful. KV-Kaizen instead keeps the input tokens and compresses how their KV information is represented.
The method combines three forms of compression. Precision stores cache values using fewer bits, such as 2, 4, or 8 bits instead of 16. Rank stores a shorter latent representation of the KV states, using multi-head latent attention and truncating that representation. Depth allows multiple layers to share a cache: a layer can inherit the most recently produced cache rather than writing its own.
Each layer makes one combined choice. It can inherit a cache, or produce one using a selected latent width and bit-width. These choices are expressed in a common cost measured in bits per token. The overall target is therefore a cache budget rather than an independent limit on each compression technique.

How the adaptive selector works

KV-Kaizen uses a small selector that reads the prompt’s token embeddings through two bidirectional transformer blocks. Its representations are averaged across the input, then separate output heads choose an action for each model layer. The selector and the backbone are fine-tuned together with a language-modeling objective and a loss that encourages the realized cache cost to match a target compression ratio.
The selector’s decision is made once, before pre-fill, and remains fixed throughout generation. This is important operationally: the system does not need to make a new decision for every decoded token, and the cache shape stays stable. In the reported implementation, selectors contain 2.6–3.7 million parameters, less than 0.05% of models with 7 billion parameters or more. On Qwen2.5-7B, selector execution takes about 3 milliseconds across input lengths from 2,000 to 32,000 tokens.
The model is trained with the cache configuration it will use at serving time. This differs from simply compressing a frozen model after training. The paper reports that applying a learned configuration to a frozen backbone causes accuracy to collapse, whereas joint fine-tuning teaches the model to operate with its compressed cache.
This context-dependent allocation contrasts with approaches that compress an entire interaction history into a shared representation. For example, Anchored Context Distillation for latent-observation software engineering agents also addresses the danger that compression can discard information needed by later actions, but it focuses on agent observations and action-preserving context layouts rather than per-layer KV-cache configurations.

What results stand out

Across instruction-following, reasoning, and long-context retrieval and question-answering evaluations, the learned selectors reached the reported Pareto frontier of accuracy versus cache size against learning-free and post-hoc baselines. On Qwen2.5-7B, a fourfold cache reduction achieved accuracy comparable to the uncompressed control. The paper reports that this fourfold reduction produced no accuracy degradation from 7B parameters upward in the evaluated models.
The authors also found that combining compression axes is more effective than applying one aggressively. At a fourfold reduction, rank alone or depth alone could hurt smaller models substantially, while the combination of all three axes was close to, or above, the accuracy of the corresponding uncompressed controls for the 7B and 14B Qwen2.5 models. The benefit comes from distributing smaller compromises across layers rather than imposing the same severe change everywhere.
The comparison with smaller models is notable. At cache reductions beyond roughly twofold in the Qwen2.5-7B experiments, a compressed 7B model scored above the uncompressed cache configurations of several smaller models at the same cache-size reference. This suggests that retaining a larger model while compressing its cache can be preferable to serving a smaller model with an ordinary cache.
KV-Kaizen can also be combined with eviction because it does not remove tokens. On a Qwen2.5-14B model at 16,000 tokens, the paper reports a fourfold reduction from cache configuration plus eviction to one position in eight, producing a 32-fold smaller decode-time cache without a significant accuracy difference. Device measurements further found that combining the approaches kept peak memory nearly flat as sequences grew, and that most tested combinations decoded faster than an uncompressed cache.

What to keep in mind

The approach requires fine-tuning the model and selector together, rather than being a universally applicable post-processing step. The experiments use a specific supervised fine-tuning recipe and evaluate IFEval, GSM8K, and RULER across models from 1.5B to 32B parameters. Results vary by model size and compression axis: at 3B, every tested configuration fell behind its uncompressed control at fourfold compression, while larger models were more tolerant.
The rank-based method also relies on converting non-MLA models to multi-head latent attention using CARE, and the paper notes that truncated latent coordinates do not automatically have the desired ordering after conversion. Fine-tuning helps account for this. Finally, the reported cache cost model excludes quantization metadata, whose storage overhead is discussed separately. Thus, the headline compression ratios describe the modeled cache representation and should be interpreted alongside implementation overheads and the accuracy costs observed for particular model sizes and configurations.

Comments (0)

No comments yet

Be the first to share your thoughts!