发表机构
Korea University; Korea Maritime and Ocean University; Konkuk University(韩国大学; 韩国海洋大学; 建国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PILAR通过实体链接断言图统一多模态文档证据,在多种智能体框架和检索后端下显著提升开放域问答的端到端性能,尤其擅长组合性、跨文档和多跳问题。
AI 中文摘要
多模态文档语料库上的开放域问答(ODQA)需要将散布在文本、表格和图形中的证据进行关联。现有系统通常将这些来源分开存储,或仅检索粗略的页面,这削弱了全局证据关联。我们提出PILAR,一种基于分页的统一证据表示,实例化为实体链接断言图。PILAR将源自句子、表格和图形的信息映射到公共断言空间,并利用该图作为稳健页面检索之上的受控关联层。在包含四种智能体框架、十四种检索后端和两个基准的共享阅读器评估中,PILAR取得了最佳端到端EM/ANLS。增益在组合性、跨文档和多模态问题上最大,单次改进比平面检索提高+1.6 EM,在组合性问题上提高到+2.9,在三跳问题上提高到+5.9。消融实验表明,当前增益主要由框架的文本实例化部分驱动,而视觉断言仅在局部感知过滤后才有帮助。因此,我们将PILAR定位为多模态ODQA的统一证据表示,而非独立的视觉推理模块。
英文摘要
Open-domain question answering (ODQA) over multimodal document corpora requires linking evidence scattered across text, tables, and figures. Existing systems often store these sources separately or retrieve only coarse pages, which weakens global evidence linking. We present PILAR, a page-grounded unified evidence representation instantiated as an entity-linked assertion graph. PILAR maps sentence-, table-, and figure-derived facts into a common assertion space and uses the graph as a controlled linking layer over robust page retrieval. In a shared-reader evaluation with four agent frameworks, fourteen retrieval backends, and two benchmarks, PILAR achieves the best end-to-end EM/ANLS. Gains are largest on compositional, cross-document, and multimodal questions, with a single-shot improvement of +1.6 EM over flat retrieval, rising to +2.9 on compositional and +5.9 on 3-hop questions. Ablations show that current gains are driven mainly by the text-instantiated slice of the framework, while visual assertions help only after locality-aware filtering. We therefore position PILAR as a unified evidence representation for multimodal ODQA rather than a standalone visual-reasoning module.
CommentsAccepted to Findings of EMNLP 2026. 25 pages, 5 figures, 24 tables