No AI summary available for this article.
Why It Matters
Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its index alone. We frame inversion as conditional document image generation and infer from the vectors...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2610.09920v1 · Indexed about 2 hours ago