arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00923cs.IR

CANOPY:多模态RAG的自适应粒度证据压缩

CANOPY: Adaptive-Granularity Evidence Compression for Multimodal RAG

Hyojeong Yun, Jueun Kim, Wook-Shin Han

首次发表
浏览论文内容

中文总结 AI 辅助

CANOPY通过层级表示和节点编码器评分实现多模态RAG的自适应粒度证据压缩,并利用评论器触发后续检索,在五个QA基准上提升准确率并减少14.2%-27.7%的输入token。

中文摘要 AI 辅助

多模态RAG检索文本、表格、图像和视频,但选择检索粒度并不能决定在每个条目内保留多少上下文。粗粒度单元包含不相关内容,而均匀细粒度选择可能移除理解证据所需的上下文。现有压缩器通过模态特定机制解决这一权衡,但缺乏一种跨异构条目逐区域调整保留范围的共享过程。我们提出CANOPY(层级规范投影),一种用于检索后自适应粒度证据压缩的框架。CANOPY将检索到的条目表示为层级结构,并使用在黄金证据上微调的节点编码器对区域与查询的相关性进行评分。父级相对细化比较这些分数,以在不同粒度下选择多个区域,无需LLM调用进行节点级剪枝。由于压缩无法恢复从未检索到的证据,评论器在判断累积证据不足时会请求有针对性的后续检索;新检索到的条目在添加前会被压缩。在包含33M条异构语料库的五个QA基准上,CANOPY实现了比评估的检索基线更高的平均答案准确率。消融实验表明,额外检索驱动了多跳QA的主要准确率提升。在未路由的Qwen3-VL-8B-Instruct设置中,与无压缩的相同迭代流程相比,压缩将读者输入证据token减少了14.2%-27.7%,同时答案准确率相当。

英文摘要

Multimodal RAG retrieves text, tables, images, and videos, but choosing a retrieval granularity does not determine how much context to retain within each item. Coarse units include irrelevant content, while uniformly fine selection can remove context needed to interpret the evidence. Existing compressors address this trade-off with modality-specific mechanisms, leaving open a shared procedure for adapting the retained extent region by region across heterogeneous items. We introduce CANOPY (Canonical Projection over Hierarchy), a framework for adaptive-granularity post-retrieval evidence compression. CANOPY represents retrieved items as hierarchies and uses a node encoder fine-tuned on gold evidence to score regions against the query. Parent-relative refinement compares these scores to select multiple regions at different granularities without LLM calls for node-level pruning. Because compression cannot recover evidence that was never retrieved, a critic requests targeted follow-up retrieval when it judges the accumulated evidence insufficient; newly retrieved items are compressed before being added. Across five QA benchmarks over a 33M-item heterogeneous corpus, CANOPY achieves higher average answer accuracy than the evaluated retrieval baselines. Ablations indicate that additional retrieval drives the main accuracy gains on multi-hop QA. In the unrouted Qwen3-VL-8B-Instruct setting, compression reduces reader-input evidence tokens by 14.2-27.7% relative to the same iterative pipeline without compression, with comparable answer accuracy.

发表机构

  • GSAI, POSTECH(全球人工智能研究所,浦项科技大学)
  • CSE, POSTECH(计算机科学系,浦项科技大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑