发表机构
The Hong Kong University of Science and Technology (Guangzhou); The University of Hong Kong; Tsinghua University; University of Sussex(香港科技大学(广州); 香港大学; 清华大学; 萨塞克斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出LedgerMind,通过结构化证据账本及三层依据协议等组件,解决多模态智能体推理中最终答案准确率无法反映轨迹可信度的问题,在多模态基准上同时提升了答案准确率与轨迹可信度。
AI 中文摘要
用于视觉问答的多模态智能体越来越多地以穿插感知、检索和推理的多步轨迹运行,但评估仍在很大程度上简化为最终答案的准确率。这种聚合信号无法区分正确答案是通过有根据的证据、语言先验还是偶然的误差抵消得出的。我们建议将多模态智能体轨迹视为一种溯源约束状态机:工具输出被标准化为结构化证据账本,作为轨迹状态;下游推理和决策主张只能引用活跃的账本条目,在实体和数值层面检查依据,修复则通过类型化状态转换实现,该转换不能引入无工具生成溯源的内容。我们将此设计实例化为LedgerMind(基于结构化证据账本的溯源约束多模态智能体推理),并补充了三层依据协议、将推理深度与问题复杂度匹配的自适应双路径调度器,以及具有正式溯源非放大保证的事件触发验证与修复引擎。我们使用LedgerMind针对最终答案准确率往往掩盖的四种反复出现的失败模式:无依据的中间推理、有引用支持的实体幻觉(幻影依据)、对简单查询的过度推理以及修复时的放大。在多个多模态推理基准和骨干多模态大语言模型(MLLM)上的实验表明,LedgerMind同时提高了答案准确率和轨迹级可信度。
英文摘要
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state machine: tool outputs are normalized into a Structured Evidence Ledger that serves as the trajectory state, downstream reasoning and decision claims may cite only active ledger entries, grounding is checked at the entity and numeric level, and repair is realized as typed state transitions that cannot introduce content without tool-produced provenance. We instantiate this design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question complexity, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. We use LedgerMind to target four recurring failure patterns that final-answer accuracy tends to obscure: unsupported intermediate reasoning, citation-backed entity hallucination (Phantom Grounding), over-reasoning on simple queries, and repair-time amplification. Experiments across multiple multimodal reasoning benchmarks and backbone MLLMs show that LedgerMind improves both answer accuracy and trajectory-level faithfulness.