arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11433cs.AI

全决策:一种用于全模态问答的渐进式证据状态智能体系统

Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents

Ming Ma, Yi Zhu, Yiran Zhong, Feida Zhu, Yuhao Wang, Junhan Shi, Lingrui Mei, Tianming Yang, Steven Hoi

首次发表
浏览论文内容

中文总结 AI 辅助

研究全模态问答中证据分散问题,提出全决策系统,将问答转化为证据闭合过程,维护结构化证据状态,通过确定性状态更新处理异构观察,提升了准确率,验证了证据状态控制在多步证据寻求中的作用。

中文摘要 AI 辅助

全模态证据寻求问答要求智能体回答证据分散在视频、音频、图像、网页和计算结果中的问题。现有多模态系统常将证据留在便签、工具轨迹或自由形式历史记录中,难以追踪已证实、缺失及足够回答的证据。我们提出全决策,一个无需训练的证据状态系统,将全模态问答转化为查询范围的证据闭合过程。对于每个查询,它维护一个结构化证据状态,包含已确认证据、未解决冲突、事实和计算依赖以及开放证据需求。共享状态视图指导规划、证据获取、验证、修复和最终确定。异构观察通过确定性状态更新进行归一化、判断和提交。该设计实现有针对性的证据获取,保留稀疏跨模态线索,并提供对修复和停止的可检查控制。全决策在OmniGAIA上达到45.6%的准确率,在WorldSense上达到58.3%,分别比基线提高27.3和30.2个百分点。无状态消融和轨迹审计进一步支持了显式证据状态控制在多步全模态证据寻求中的作用。

英文摘要

Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions. Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend. We present Omni-Decision, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict. A critic reads each noisy observation and passes only the usable content to the ledger, discarding the rest, so the planner works from a compact context throughout the task. Each run records the state, action, and verdict at every step, and supervised fine-tuning and decision-level reinforcement learning on these trajectories further improve the planner. Omni-Decision achieves state-of-the-art accuracy of 81.4% on OmniGAIA at approximately 43% of Gemini-3.1-Pro's cost per question, and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model.

发表机构

  • Institute of Neuroscience, Chinese Academy of Sciences(中国科学院神经科学研究所)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室)
  • Tsinghua University(清华大学)
  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑