发表机构
National University of Singapore; Fudan University; Tencent(新加坡国立大学; 复旦大学; 腾讯)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估全模态生成模型 MiniMax-H3 的物理世界推理能力,通过四类跨模态任务(517 实例)发现其总体成功率 41.97%,表明多模态整合是提升推理的关键。
AI 中文摘要
近期,全模态生成模型(Omni-Modal Generative Models,简称 Omni-Models)已将内容生成推进到文本、图像、视频和音频的统一建模。MiniMax-H3 通过在共享潜在框架中结合多模态上下文理解与联合音视频生成,体现了这一转变。其统一架构提出了一个基本问题:多模态对齐能否提升模型的世界推理能力,以及全模态输入能够启用哪些新的评估范式?为探究这一问题,本工作引入了一个围绕物理世界推理的四个互补维度组织的综合评估框架。与现有的视频生成和世界模型评估框架不同——后者通常受限于有限的输入模态以及提示与目标视频内容高度匹配的评估设置——我们的评估专门设计用于利用 Omni-Model 的多模态输入。我们构建了多样化的新任务集合,要求模型整合跨模态的互补信息。具体而言,我们考虑了四种场景,包括隐式提示配多帧、音频-图像、前缀视频以及音频-视频输入。每种单一模态仅提供关于底层事件的部分证据,要求模型对互补语义线索进行联合推理,以推断潜在事件状态和未来动态。在 517 个评估实例中,MiniMax-H3 实现了 41.97% 的总体成功率。基于视频的决策推理成功率最高,达 56.00%,而基于音频的去歧义推理最弱,仅为 27.40%。这些结果表明,有效的多模态整合仍是充分利用多样化输入模态优势的关键。项目可在以下网址获取:此 https URL。
英文摘要
Recent Omni-Modal Generative Models (Omni-Models) have advanced content generation toward unified modeling of text, images, video, and audio. MiniMax-H3 exemplifies this transition by combining multimodal context understanding with joint audio-visual generation in a shared latent framework. Its unified architecture raises a fundamental question: Can multimodal alignment improve the model's world reasoning, and what new evaluation paradigms do omni-modal inputs enable? To investigate this question, this work introduces a comprehensive evaluation framework organized around four complementary dimensions of physical world reasoning. Unlike existing evaluation frameworks for video generation and world models, which are often constrained by limited input modalities and evaluation settings where prompts closely match the target video content, our evaluation is specifically designed to exploit the multimodal inputs of Omni-Model. We construct a diverse set of novel tasks that require models to integrate complementary information across modalities. Specifically, we consider four scenarios, including implicit prompts paired with multiple frames, audio-image, prefix-videos, and audio-video inputs. Every single modality provides only partial evidence about the underlying event, requiring the model to jointly reason over the complementary semantic cues to infer latent event states and future dynamics. Across 517 evaluation instances, MiniMax-H3 achieves an overall success rate of 41.97%. Video-based Decision Reasoning yields the highest success rate at 56.00%, while Audio-based Disambiguation Reasoning is the weakest, reaching only 27.40%. These results indicate that effective multimodal integration remains key to fully exploiting the benefits of diverse input modalities. The project is available at https://github.com/gulucaptain/MiniMax-H3-Reason.
Comments17 pages, 14 figures