Inverting Multi-Vector Visual Document Indices
Researchers invert multi-vector visual document indices, recovering 47% of words and 45% of sensitive tokens on ViDoRe v3.
The paper shows multi-vector visual document retrievers, which store about a thousand patch vectors per page, can be inverted to reconstruct a page from its index alone. On ViDoRe v3, inversions from raw indices recover 47% of words and 45% of sensitive tokens, and as queries rank the source page first 98.4% of the time. Token pooling and shuffling cut word recall to about 8%, but a model that restores shuffled order raises first-rank recovery from 3.8% to 93.5%. Applied unchanged to another multi-vector retriever, inverted pages still rank their source first 70.2% of the time.
- Raw-index inversion recovers 47% of words and 45% of sensitive tokens
- Inverted pages rank their source page first 98.4% of the time
- Pooling and shuffling cut word recall to about 8%
- Order restoration lifts shuffled first-rank rate from 3.8% to 93.5%
- The same attack ranks source pages first 70.2% on another retriever
Full article255 words · extracted from arxiv.org · click to collapse
Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its index alone. We frame inversion as conditional document image generation and infer from the vectors what the attack needs: the encoder, the page shape and, for shuffled vectors, their order. On the ViDoRe v3 benchmark, pages inverted from raw indices recover 47% of the words and 45% of the sensitive tokens. Used as queries against the stored indices, they rank their source page first 98.4% of the time. We test two cheap protections, token pooling and shuffling, which both cut word recall to about 8%. A model that restores the order of a shuffled index raises the share of source pages ranked first from 3.8% to 93.5%, while inverting a pooled index remains open. To test generalisation, we apply the same attack unchanged to another multi-vector retriever: its inverted pages still rank their source page first 70.2% of the time, though its word recall stays below a nearest-neighbour baseline. Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.09920