arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.08732cs.CVcs.CLcs.IR

AnchorFold:基于递归注意力传播的“先聚焦后折叠”高效多向量视觉文档检索框架

AnchorFold: A Focus-Then-Fold Framework via Recursive Attention Propagation for Efficient Multi-Vector Visual Document Retrieval

Haoyu Zuo, Yibo Yan, Xin Zou, Shuliang Liu, Yi Cao, Mingdong Ou, Xuming Hu

AI总结:

AnchorFold是无训练的文档侧索引压缩框架,通过递归注意力传播实现先聚焦后折叠,在多数据集多骨干下,强压缩时仍保持高检索性能,优于同类无训练基线。

AI中文摘要:

多向量视觉-语言检索器通过后期交互实现细粒度视觉文档检索(VDR),但每页存储和评分数百个视觉补丁嵌入会产生大量开销。现有无训练方法依赖剪枝或合并:剪枝在强压缩下性能大幅下降,合并在生成代表时未明确优先重要区域。本文提出AnchorFold,一种用于文档侧索引压缩的无训练“先聚焦后折叠”框架。AnchorFold在视觉自注意力图上应用递归注意力传播,在每个注意力头内执行多步传播并整合各头和层的分数。聚焦阶段选择中心性最高的token作为锚点;折叠阶段在归一化检索空间中将剩余token分配给最相似的锚点,通过中心性加权聚合总结每个锚点为中心的组,在保留非锚点贡献的同时将容量集中于结构重要的token。在ViDoRe v1/v2和REAL-MM-RAG数据集上使用三种不同检索骨干的实验显示,当压缩比γ≤0.20时,AnchorFold始终优于所有评估的无训练基线;在ViDoRe v1/v2上,5倍压缩时平均保留完整索引98.3%的NDCG@5,实现近无损压缩,20倍压缩时保留92.4%。

英文摘要:

Multi-vector vision-language retrievers enable fine-grained Visual Document Retrieval (VDR) through late interaction, but storing and scoring hundreds of visual patch embeddings per page incurs substantial overhead. Existing training-free methods rely on pruning or merging: pruning degrades sharply under aggressive compression, whereas merging does not explicitly prioritize important regions when forming representatives. We introduce AnchorFold, a training-free focus-then-fold framework for document-side index compression. AnchorFold applies Recursive Attention Propagation over visual self-attention graphs, performing multi-step propagation within each attention head and integrating scores across heads and layers. The focus stage selects the highest-centrality tokens as anchors. The fold stage assigns remaining tokens to their most similar anchors in the normalized retrieval space and summarizes each anchor-centered group through centrality-weighted aggregation. This preserves non-anchor contributions while concentrating capacity on structurally important tokens. Across ViDoRe v1/v2 and REAL-MM-RAG with three diverse retrieval backbones, AnchorFold consistently outperforms all evaluated training-free baselines at $γ\leq 0.20$. On ViDoRe v1/v2, it retains 98.3% of full-index NDCG@5 on average at $5\times$ compression, achieving near-lossless compression, and 92.4% at $20\times$ compression.

补充信息

↑