Walking the Embedding Space: Datastore Extraction from Multimodal RAG
Researchers show a black-box attack can extract hundreds of private images from multimodal RAG datastores.
The paper presents a black-box extraction attack against image-returning multimodal retrieval-augmented generation, where the retrieved image itself is the response. Each query blends an attacker-held shadow image with an already recovered image, and relevance-weighted resampling steers later queries toward embedding regions that still yield new items. Unlike prompt-based leakage attacks, the malicious instructions are placed inside the input image. Across medical, document, and general-purpose settings, one 2,500-query run reconstructed up to 611 radiology images, 566 document scans, and 416 general images, as much as 5.6 times a non-adaptive baseline, using CLIP-family retrievers.