arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14841cs.AI

重排序器所见:面向长文档多模态问答的多方面页面标注

What the Reranker Sees: Multi-Aspect Page Annotation for Long-Document Multimodal Question Answering

  • Emory University(埃默里大学)
  • Hippocratic AI(希波克拉底人工智能公司)

机构由 AI 辅助整理,请以论文原文为准。

Guanchen Wu, Jiayuan Ding, Subhabrata Mukherjee, Carl Yang

AI总结:

本文针对长文档多模态问答的重排序瓶颈,提出含Trident-R与Trident-S组件的Trident模型,通过多方面页面标注提升检索与生成性能,在多数据集上取得显著效果。

AI中文摘要:

针对数十至数百页混合文本、表格、图表与图像的长文档视觉问答(VQA),通常采用检索-后读取的流程。在本文设定中,瓶颈从检索召回率转向重排序器侧的证据选择:在MMLongBench-Doc数据集上,BGE-M3的Recall@20达0.86,但F1@5仅为0.254;即便视觉检索器ColPali,其F1@5也仅达0.332;仅查看原始片段的纯文本重排序LLM,即便上游检索器已编码图像,仍会遗漏表格、图表及布局证据。本文提出Trident模型,包含两个互补组件:Trident-R是与检索器无关的LLM重排序器,它将每个候选转换为LLM可读的语义记录,包含视觉描述、章节路径、实体标签、多轴概念命中及文本片段,随后执行单次自适应K重排序调用;Trident-S是生成侧模块,在合成前从主题、实体及结构视角提示VLM。在两个长文档数据集上,标注+重排序协议显著提升了五个异构池的检索F1,每个重排序池均超过最强的自适应K基线PageIndex。无标注的LLM重排序几乎未改变首命中排名,表明性能提升源于结构化标注。Trident-S专为开放式合成问题设计,使这类问题的生成准确率提升最多6.6个百分点。最优Trident配置是本次评估中最强的下游QA流程,在两个LLM评判者间排名一致(kappa=0.913)。

英文摘要:

Long-document visual question answering (VQA) over documents of tens to hundreds of pages mixing text, tables, charts, and figures typically follows retrieve-then-read pipelines. In our setting, the bottleneck shifts from retrieval recall to reranker-side evidence selection: on MMLongBench-Doc, BGE-M3 reaches Recall@20 = 0.86 but only F1@5 = 0.254, and even the visual retriever ColPali reaches only F1@5 = 0.332; a text-only rerank LLM seeing only raw snippets misses table, chart, and layout evidence even when the upstream retriever encoded images. We propose Trident, with two complementary components: Trident-R, a retriever-agnostic LLM reranker that converts each candidate into an LLM-readable semantic record, including a visual caption, section path, entity tags, multi-axis concept hits, and a text snippet, then performs a single adaptive-K rerank call; and Trident-S, a generation-side module that prompts the VLM under topical, entity, and structural lenses before synthesis. On two long-document datasets, the annotation+rerank protocol substantially improves retrieval F1 across five heterogeneous pools, with every reranked pool exceeding the strongest adaptive-K baseline PageIndex. An LLM rerank without the annotation barely changes first-hit ranking, indicating the lift comes from the structured annotation. Trident-S targets open-ended synthesis questions by design, adding up to 6.6 points in generation accuracy on these questions. The best Trident configuration is the strongest downstream QA pipeline in our evaluation, with rankings consistent across two LLM judges (kappa = 0.913).

↑