MRVQ: One Resident Index for Dimension- and Rate-Elastic Vector Search
MRVQ truncates one residual quantization code by rate and dimension, using 17.8-22x less memory than separate QINCo2 indices for elastic dense retrieval.
Matryoshka Residual Vector Quantization (MRVQ) is a post-hoc residual quantizer for frozen embeddings whose code truncates by dropping residual stages (rate) or coordinates (dimension), serving every (dimension, rate) pair from one artifact. Across FiQA and NFCorpus, four embedding families, and 4/8/16-byte codes, it uses 17.8-22.0x less memory than three separately trained QINCo2 indices and 1.89-2.02x less than a shared-model baseline. Quality costs 0.026-0.107 nDCG@10 versus per-rate QINCo2 on FiQA, but MRVQ beats PQ, OPQ, and AdANNS-OPQ at matched code size. A low-build-cost PCA-scalar design matches RaBitQ-quality while fitting 420x faster at the median.
- Single resident code stream serves every dimension and rate combination
- 17.8-22.0x less RAM than three separately trained QINCo2 indices
- Quality cost of 0.026-0.107 nDCG@10 versus per-rate QINCo2 on FiQA
- PCA-scalar variant builds 420x faster at median with RaBitQ-comparable quality
Full article211 words · extracted from arxiv.org · click to collapse
Dense-retrieval services must switch among embedding-prefix dimensions and index bit rates as latency, quality, and memory budgets change. Tuning a quantizer separately for each rate gives the best quality, but the retrieval tier then holds several code streams and quantizer states at once. We introduce Matryoshka Residual Vector Quantization (MRVQ), a post-hoc residual quantizer for frozen embeddings. Its maximum-rate code can be truncated two ways: dropping residual stages lowers the rate, and dropping embedding coordinates lowers the dimension. One resident artifact therefore serves every (dimension, rate) pair we evaluate. Across FiQA and NFCorpus, four embedding families, and {4, 8, 16}-byte codes, MRVQ is the lowest-RAM design we evaluate. It uses 17.8-22.0x less memory than three separately trained QINCo2 indices, and 1.89-2.02x less than a lean shared-model steelman. The saving is not free: per-rate QINCo2 is 0.026-0.107 nDCG@10 better on FiQA. But MRVQ beats PQ, OPQ, and AdANNS-OPQ at matched code size. We also evaluate a low-build-cost PCA-scalar design that attains quality comparable to RaBitQ and its extension while fitting 420x faster at the median. Finally, we report two negative results: QINCo2 collapses when trained at high rates, and a ranking-bound hypothesis misses its pre-specified acceptance criteria. MRVQ is therefore a low-memory operating point for elastic retrieval, not a universal quality winner.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.03651