AI 中文总结
针对长音频-视频推理难题,OmniReasoner提出工具使用后训练框架,通过监督微调与强化学习让全模态语言模型学会调用放大工具,利用TimeAnchor保持时间参数一致性,借助时间增强数据引擎实现可训练,提升了答案准确性与时间定位能力。
AI 中文摘要
长音频-视频推理对全模态语言模型来说具有挑战性,因为决定性证据往往稀疏、跨模态且难以以统一的高保真输入保存。我们引入了OmniReasoner,这是一个用于长音频-视频推理的工具使用后训练框架。全模态语言模型通过监督微调与强化学习,学会在回答前决定是否以及在何处调用放大工具。OmniReasoner先构建低成本的全局预览,必要时调用放大工具进行高保真视听检查。我们还引入TimeAnchor来保持工具时间参数的有效性与一致性。为使工具使用行为可训练,构建了时间增强数据引擎。实验表明,OmniReasoner提高了答案准确性与时间定位能力。
英文摘要
Long audio-video reasoning is difficult for omnimodal LLMs because the decisive evidence is often sparse, cross-modal, and too expensive to preserve with uniformly high-fidelity inputs. We introduce OmniReasoner, a tool-use post-training framework for Thinking with Long Audio-Video: omni-modal LLMs learn, via supervised fine-tuning and reinforcement learning, to decide whether and where to call a zoom-in tool before answering. OmniReasoner first builds a low-cost global preview of the full stream and then, when needed, calls the zoom-in tool with a requested temporal interval for higher-fidelity visual and audio inspection before answering. Because the model observes different sampling granularities before and after this call -- a sparse global preview and a denser local clip -- we introduce TimeAnchor, which keeps the tool's temporal argument valid and round-trip-consistent across these granularities, rather than tied to frame indices from a particular sampling rate. To make this tool-use behavior trainable without expensive manual interval annotation, we build a Temporal Augmented Data Engine that synthesizes tool-use post-training trajectories by video editing and composition. Experiments across omnimodal and video benchmarks show that OmniReasoner improves both answer accuracy and temporal grounding while concentrating high-fidelity computation on informative regions. Code is available at https://github.com/RockyChen0205/OmniReasoner.