arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25738cs.AI

OmniFysics-Nano-V2 技术报告:跨模态理解物理世界

OmniFysics-Nano-V2 Technical Report: Understanding the Physical World Across Modalities

Yizhou Liu, Jinghang Han, Kaixiang Qiu, Qi He, Minghao Han, Yue Jiang, Xujia Chen, Wei Zou, Shunli Wang, Lihua Zhang, Dingkang Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对全模态模型缺乏物理监督与同质训练目标的问题,提出OmniFysics-Nano-V2,采用双分支物理感知数据流水线与两阶段GRPO课程,在21个基准中17个取得领先,提升物理世界理解。

中文摘要 AI 辅助

全模态模型已将多模态交互扩展到视觉、音频、语音和语言领域。然而,它们的训练主要围绕语义描述和通用目标进行组织,导致物理属性、交互状态和因果机制仅被部分指定。这一差距不仅仅是模态覆盖范围的问题:增加更多模态本身并不能提供将观察结果与世界的物理结构联系起来的监督信号。我们提出了 OmniFysics-Nano-V2,一个用于物理世界感知与理解的紧凑型全模态模型。该模型在共享推理框架内支持图像、视频、音频、语音和文本输入,并支持文本和语音生成。为了解决缺乏显式物理监督的问题,我们构建了一个双分支物理感知数据流水线,将显著对象锚定在结构化物理属性中,并将视觉变化与声学事件、中间响应和交互结果对齐。为了解决同质化训练目标的问题,我们根据奖励多样性策划强化学习提示,并采用两阶段群体相对策略优化课程,从一般任务正确性逐步过渡到细粒度物理感知推理。跨多模态、音视频和物理推理基准的实验表明,所提出的数据和训练策略在保持广泛全模态能力的同时,提升了物理世界理解能力。所提出的模型在 21 个基准中的 17 个上取得了领先结果,优于最先进的全模态模型。通过为 AI 系统配备全模态和物理世界感知能力,OmniFysics-Nano-V2 有望成为下一代物理 AI 的基石。

英文摘要

Omni-modal models have expanded multimodal interaction across vision, audio, speech, and language. However, their training is predominantly organized around semantic descriptions and general-purpose objectives, leaving physical attributes, interaction states, and causal mechanisms only partially specified. This gap is not simply a matter of modality coverage: adding more modalities does not by itself provide the supervision needed to connect observations with the physical structure of the world. We present OmniFysics-Nano-V2, a compact omni-modal model for physical-world perception and understanding. The model supports image, video, audio, speech, and text inputs within a shared reasoning framework, together with text and speech generation. To address the lack of explicit physical supervision, we construct a dual-branch physics-aware data pipeline that grounds salient objects in structured physical attributes and aligns visual changes with acoustic events, intermediate responses, and interaction outcomes. To address homogeneous training objectives, we curate reinforcement-learning prompts by reward diversity and adopt a two-stage Group Relative Policy Optimization curriculum that progresses from general task correctness to fine-grained physical perceptual reasoning. Experiments across multimodal, audio-visual, and physical reasoning benchmarks show that the proposed data and training strategy improves physical-world understanding while preserving broad omni-modal competence. The proposed model achieves leading result on 17 of 21 benchmarks against SOTA omni-modal models. By equipping AI systems with both omni-modal and physical-world perception capabilities, OmniFysics-Nano-V2 is poised to become a cornerstone of next-generation Physical AI.

发表机构

  • Physical Superintelligence Lab, Fysics AI(Fysics AI 物理超级智能实验室)
  • College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑