arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15128cs.LG

流式全模态思考

Omni-Streaming Thinking

Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li, Yiwen Guo, Yongqi Zhang, Difan Zou

首次发表
浏览论文内容

中文总结 AI 辅助

针对流式全模态模型过早跨模态承诺问题,提出Omni-Streaming Thinking方法,通过结构化证据、预测与验证机制及答案门控,在五个基准上平均相对提升超10%,并引入诊断基准OST-DiagBench验证其有效性。

中文摘要 AI 辅助

流式全模态模型必须从迄今观察到的视频片段和同步音频中决定回答什么以及何时回答。视觉线索往往在话语或声音事件完成之前就支持某种解释。如果该解释作为事实进入记忆,后续推理可能会持续依赖它,即使音频后来与之矛盾。我们将这种失败称为过早跨模态承诺。我们提出Omni-Streaming Thinking (OST),它生成结构化输出,包括迄今观察到的证据、对未来证据的预测,以及基于这些证据的主张。每个主张最初被标记为待定,并链接到未来的验证区间。音频和视觉证据被分开存储,OST在验证区间结束时根据指定模态的证据检查主张。当检测到矛盾证据时,反驳过程会降低该主张及其依赖状态的影响,然后使用新证据指导状态更新。答案门控决定答案关键主张是否满足给出响应的条件。使用冻结的Qwen3-Omni-30B-A3B-Instruct骨干网络配合轻量级适配,OST在五个流式和音视频基准上平均相对性能比最强的开放基线高出超过10%。我们还引入了OST-DiagBench,它固定视频并编辑音频以测试一致性、缺失、矛盾、共存和字幕-语音冲突。OST达到d-prime = 2.95,而开放基线最多为1.38,同时减少了视觉引发的听觉幻觉。

英文摘要

Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If that interpretation enters memory as a fact, later reasoning can keep relaying it even after audio contradicts it. We call this failure premature cross-modal commitment. We propose Omni-Streaming Thinking (OST), which generates structured outputs that include evidence observed so far, forecasts of future evidence, and claims based on this evidence. Each claim is initially marked as pending and linked to a future verification interval. Audio and visual evidence are stored separately, and OST checks a claim against the evidence from the specified modality at the end of the verification interval. When contradictory evidence is detected, a refutation process reduces the influence of the claim and its dependent states, and then guides a state update using the new evidence. An answer gate decides whether the answer-critical claims meet the conditions for giving a response. Using a frozen Qwen3-Omni-30B-A3B-Instruct backbone with lightweight adaptation, OST outperforms the strongest open baselines on five streaming and audio-visual benchmarks by more than 10% relative on average. We also introduce OST-DiagBench, which holds video fixed and edits audio to test agreement, absence, contradiction, coexistence, and subtitle-speech conflict. OST reaches d-prime = 2.95, compared with at most 1.38 for open baselines, while reducing vision-induced auditory hallucinations.

发表机构

  • The University of Hong Kong(香港大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • University of Sussex(萨塞克斯大学)
  • LIGHTSPEED

机构由 AI 辅助整理,请以论文原文为准。

相关深度报道

↑