Inverting Multi-Vector Visual Document Indices
Researchers invert multi-vector visual document indices, recovering 47% of words and 45% of sensitive tokens on ViDoRe v3.
The paper shows multi-vector visual document retrievers, which store about a thousand patch vectors per page, can be inverted to reconstruct a page from its index alone. On ViDoRe v3, inversions from raw indices recover 47% of words and 45% of sensitive tokens, and as queries rank the source page first 98.4% of the time. Token pooling and shuffling cut word recall to about 8%, but a model that restores shuffled order raises first-rank recovery from 3.8% to 93.5%. Applied unchanged to another multi-vector retriever, inverted pages still rank their source first 70.2% of the time.