Back to AI Research

AI Research

MRVQ compresses a flexible vector-search menu into one resident index

Key Takeaways

  • The study reports smaller resident index state at a measurable retrieval-quality cost, with no end-to-end latency claim.
  • A retrieval service may need several settings for embedding dimension and compressed document size.
  • Keeping a separately optimized index for each setting consumes memory even when only one setting is serving a query.
  • Sean Culatana, Shang-En Huang and Kang Li propose MRVQ, a residual quantizer whose stored codes support both kinds of adjustment.
  • The [MRVQ paper](https://arxiv.org/abs/2610.03651) describes a specific deployment trade-off.

A retrieval service may need several settings for embedding dimension and compressed document size. Keeping a separately optimized index for each setting consumes memory even when only one setting is serving a query. Sean Culatana, Shang-En Huang and Kang Li propose MRVQ, a residual quantizer whose stored codes support both kinds of adjustment.
The MRVQ paper describes a specific deployment trade-off. One resident artifact covers every dimension-and-rate pair evaluated, while separately trained QINCo2 indices achieve better retrieval quality on FiQA. The authors report both outcomes instead of describing compression as an unconditional improvement.

One stored code, two ways to shorten it

MRVQ operates after an embedding model has produced frozen document vectors. Each residual stage contributes a one-byte code index. Reading fewer stages lowers the number of bytes used per vector; reading fewer embedding coordinates lowers the reconstruction dimension.
The evaluated serving menu uses four-, eight- and sixteen-byte codes. A service can select a shorter stage prefix without storing another independently encoded document stream. The fitting objective accounts for multiple coordinate prefixes so that dimension changes are part of the same resident design.
The paper also separates document-code storage from fixed quantizer state. For a three-rate menu, the separate document streams total 28 bytes per vector, compared with one 16-byte stream. The larger reported memory ratios at small corpus sizes depend heavily on fixed model-state overhead, not only document compression.

Memory savings have a measured quality cost

The quality study covers FiQA and NFCorpus with four embedding families: MPNet, Mxbai, Nomic and BGE. FiQA contains 57,638 documents; NFCorpus contains 3,633. The authors evaluate ranking quality with nDCG@10 and use paired per-query bootstrap intervals.
Across the evaluated configurations, MRVQ uses 17.8–22.0 times less memory than three separately trained QINCo2 indices. Against a lean shared-model alternative, the ratio is 1.89–2.02. These totals count parameter and code bytes analytically, excluding runtime buffers and optimizer state; they are not measured process memory readings.
Separately trained QINCo2 is 0.026–0.107 nDCG@10 better on FiQA at matched code size. MRVQ outperforms the evaluated PQ, OPQ and AdANNS-OPQ baselines at matched bytes, placing it between simpler compressors and the higher-quality specialized alternative.

Limits for larger retrieval services

The advantage changes with corpus size. As document storage dominates fixed quantizer state, the separate-index ratio approaches the 28-to-16 byte ratio, or 1.75. The shared-model advantage narrows further because both designs store the same maximum-rate bytes per vector. A small per-tenant index and a very large global index therefore face different memory arithmetic.
The study also reports a failed ranking-bound acceptance test and high-rate QINCo2 training collapses in some runs. Those failures remain in the analysis rather than disappearing from the comparison.
The authors measure ranking quality and resident-state accounting, not end-to-end search latency or throughput. MRVQ gives memory-constrained retrieval teams an operating point to investigate. Its memory ratios should not become speedup claims, and the results on two modest corpora leave wider deployment validation open.

Comments