Multi-vector visual document indices can leak page text
A paper shows multi-vector visual document indices can be inverted, recovering 47% of words and 45% of sensitive tokens on ViDoRe v3.
Two reports describe a paper showing that multi-vector visual document retrievers, which store about a thousand patch vectors per page, can be inverted to reconstruct page images and recover text. Inversion is framed as conditional document image generation; one source says it requires the encoder, page shape, and vector order, while the other says a page can be reconstructed from its index alone. On ViDoRe v3, inversions from raw indices recover 47% of words and 45% of sensitive tokens, and as queries rank the source page first 98.4% of the time. Token pooling and shuffling cut word recall to about 8%, but an order-restoration model raises shuffled first-rank matches from 3.8% to 93.5%. Applied unchanged to another multi-vector retriever, inverted pages still rank their source first 70.2% of the time.
- Multi-vector visual document retrievers store about a thousand patch vectors per page.
- On ViDoRe v3, inversions from raw indices recover 47% of words and 45% of sensitive tokens.
- Inverted pages used as queries rank the source page first 98.4% of the time.
- Token pooling and shuffling cut word recall to about 8%.
- An order-restoration model raises shuffled first-rank source matches from 3.8% to 93.5%.
- The same attack on another multi-vector retriever ranks the source page first 70.2% of the time.
- One report says inversion requires the encoder, page shape, and vector order; the other says a page can be reconstructed from its index alone.
Coverage timelineoldest first · each row is one article
- · 2d agoInverting Multi-Vector Visual Document Indices
Hugging Face daily papers· 66
Stored multi-vector document indices can be inverted to recover page words and sensitive tokens.
- · 1d agoInverting Multi-Vector Visual Document Indices
arXiv cs.AI / cs.LG / cs.CL· 61
Researchers invert multi-vector visual document indices, recovering 47% of words and 45% of sensitive tokens on ViDoRe v3.