OmniReasoning:推动音视频联合推理的极限
OmniReasoning: Pushing the Limits of Audio-Visual Joint Reasoning
浏览论文内容
中文总结 AI 辅助
针对音视频联合推理评估不足与能力未激发的问题,提出基准OmniReasoningBench、数据引擎OmniQA及模态因子自蒸馏方法,显著提升模型性能。
中文摘要 AI 辅助
近期进展使得统一的全模态模型能够理解音频、视觉和语言。然而,现有基准、训练数据和学习方法大多将模态独立处理,导致音视频联合推理的能力评估不足且未被充分激发。我们通过一个基准、数据引擎和学习方法来解决这一空白。首先,我们引入OmniReasoningBench,这是一个音频和视觉证据均不可或缺的基准。它包含1,150道多项选择和开放式问题,涵盖两个任务:视频内推理和视频外推理。其次,我们开发了数据引擎OmniQA。它自动构建需要明确进行音视频联合推理的基于证据的问答对,并附带时间戳线索链以指导思维过程的标注。除基准外,该引擎还生成训练数据OmniReasoning-SFT-112K和OmniReasoning-RL-19K。最后,我们提出一种在线策略自蒸馏方法——模态因子自蒸馏(MFSD)。它在模态特定的线索上下文中评估每个采样响应,解耦单个线索及其跨模态交互对令牌级信用分配的贡献。利用我们的训练数据和学习方法,我们的模型OmniReasoning-30B-A3B在OmniVideoBench上达到50.0%,在OmniReasoningBench上达到42.5%,分别比基础模型Qwen3-Omni-30B-A3B-Thinking提高了12.8和9.3个百分点。此外,它在通用和长视频基准(包括Video-MME-v2)上也取得了显著提升。我们希望我们的工作能为促进全模态联合推理的未来研究提供坚实一步。
英文摘要
Recent advances have enabled unified omni-modal models in understanding audio, vision, and language. However, existing benchmarks, training data, and learning methods largely treat the modalities independently, leaving the capability of audio-visual joint reasoning poorly evaluated and insufficiently elicited. We address this gap with a benchmark, data engine, and learning method. First, we introduce OmniReasoningBench, a benchmark where both audio and visual evidence are indispensable. It comprises 1,150 multiple-choice and open-ended questions across two tasks, reasoning over video and reasoning beyond video. Second, we develop a data engine OmniQA. It automatically constructs evidence-grounded QA pairs that explicitly necessitate audio-visual joint reasoning, together with time-stamped clue chains that guide the annotation of thinking process. Besides our benchmark, this engine produces training data OmniReasoning-SFT-112K and OmniReasoning-RL-19K. Finally, we propose an on-policy self-distillation method Modality-Factored Self-Distillation (MFSD). It evaluates each sampled response under modality-specific clue contexts, disentangling the contributions of individual clues and their cross-modal interactions for token-level credit assignment. With our training data and learning method, our model OmniReasoning-30B-A3B achieves 50.0% on OmniVideoBench and 42.5% on OmniReasoningBench, improving the base model Qwen3-Omni-30B-A3B-Thinking by 12.8 and 9.3 percentage points, respectively. Moreover, it delivers substantial gains on general and long-video benchmarks, including Video-MME-v2. We hope our work offers a solid step for facilitating future research in omni-modal joint reasoning.
发表机构
- Peking University(北京大学)
- Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。