arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多向量视觉文档索引的反演

Inverting Multi-Vector Visual Document Indices

Zhuchenyang Liu, Yao Zhang, Yu Xiao

arXiv 2610.09920首次发表:更新:

发表机构

Aalto University(阿尔托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多向量视觉文档检索器,提出从存储索引反演原始页面的攻击方法,在ViDoRe v3上恢复47%单词,并验证了池化与打乱等防护措施的局限性。

AI 中文摘要

主流的多向量视觉文档检索器将每一页存储为大约一千个补丁向量,通常存储在由第三方运营的向量数据库中。由于没有人能从向量中读取页面内容,这种索引容易被认为不如页面本身敏感。然而,由于索引按光栅顺序为每个补丁保留一个向量,且每个向量由预训练用于阅读文档的视觉语言模型计算,我们假设任何运营或入侵存储库的人都能仅凭索引重现页面。我们将反演问题构建为条件文档图像生成,并从向量中推断攻击所需的信息:编码器、页面形状,以及对于被打乱的向量,它们的顺序。在ViDoRe v3基准上,从原始索引反演的页面恢复了47%的单词和45%的敏感令牌。当用作查询对存储的索引进行检索时,它们有98.4%的时间将源页面排在首位。我们测试了两种廉价的保护措施,即令牌池化和打乱,两者都将单词召回率降至约8%。一个恢复打乱索引顺序的模型将源页面排在首位的比例从3.8%提升到93.5%,而对池化索引的反演仍然可行。为了测试泛化能力,我们将相同的攻击原封不动地应用于另一个多向量检索器:其反演页面仍有70.2%的时间将源页面排在首位,尽管其单词召回率仍低于最近邻基线。因此,多向量视觉文档检索器容易通过其存储的索引受到反演攻击,该索引应像其编码的文档一样受到保护。

英文摘要

Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its index alone. We frame inversion as conditional document image generation and infer from the vectors what the attack needs: the encoder, the page shape and, for shuffled vectors, their order. On the ViDoRe v3 benchmark, pages inverted from raw indices recover 47% of the words and 45% of the sensitive tokens. Used as queries against the stored indices, they rank their source page first 98.4% of the time. We test two cheap protections, token pooling and shuffling, which both cut word recall to about 8%. A model that restores the order of a shuffled index raises the share of source pages ranked first from 3.8% to 93.5%, while inverting a pooled index remains open. To test generalisation, we apply the same attack unchanged to another multi-vector retriever: its inverted pages still rank their source page first 70.2% of the time, though its word recall stays below a nearest-neighbour baseline. Multi-vector visual document retrievers are therefore vulnerable to inversion through their stored index, which should be protected like the documents it encodes.

Comments30 pages. Under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑