SAGE:面向中国古代文献理解的从直接回答到证据导向推理
SAGE: From Direct Answering to Evidence-Grounded Inference for Chinese Ancient Document Understanding
- Fudan University(复旦大学)
- University of Liverpool(利物浦大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有大型视觉语言模型(LVLMs)在古代文献理解中证据支撑不足的问题,本文提出SAGE多智能体框架,通过多阶段证据导向推理实现任务规划、证据获取与验证,在AncientDoc基准上优于基线,搭载Qwen3.5-9B的SAGE性能超更大单体模型。
AI中文摘要:
中国古代文献理解需要复杂的视觉、语言及历史推理。当前大型视觉语言模型(LVLMs)通常依赖不透明的单步生成范式,常产生过度自信且证据支撑不足的响应。为解决该问题,本文提出SAGE,这是一个证据导向的多智能体框架,将中国古代文献理解重新定义为证据导向推理而非直接生成答案。SAGE在受限共享状态运行时下协调专用智能体,完成任务感知规划、工具介导的证据获取、声明级验证及受限重规划。该设计支持受限证据搜索、答案修正,以及在证据不足时弃权(不执行)。在AncientDoc基准上的实验表明,SAGE在三种LVLM主干模型上均持续优于匹配的直接回答基线。值得注意的是,搭载Qwen3.5-9B的SAGE在多数评估指标上超越了大得多的单体LVLMs,凸显了超越模型规模的结构化证据导向推理的重要性。
英文摘要:
Chinese ancient document understanding demands complex visual, linguistic, and historical reasoning. Current Large Vision-Language Models (LVLMs) typically rely on an opaque, single-pass generation paradigm, often producing overconfident and weakly grounded responses. To address this, we propose SAGE, an evidence-grounded multi-agent framework that reformulates Chinese ancient document understanding as evidence-grounded inference rather than direct answer generation. SAGE coordinates specialized agents for task-aware planning, tool-mediated evidence acquisition, claim-level verification, and bounded replanning under a constrained shared-state runtime. This design supports bounded evidence seeking, answer revision, and abstention when grounding is insufficient. Experiments on the AncientDoc benchmark show that SAGE consistently outperforms matched direct-answering baselines across three LVLM backbones. Remarkably, SAGE with Qwen3.5-9B surpasses much larger monolithic LVLMs on most evaluated metrics, highlighting the importance of structured, evidence-grounded inference beyond model scaling.