AI 中文总结
针对现有全模态描述生成器忽视物理证据的问题,提出统一框架,通过数据构建、训练和评估,利用物理感知模型提取线索,训练出在多个基准上达到最先进水平的描述生成器。
AI 中文摘要
构建具有物理智能的全模态模型需要细粒度的监督信号,以捕获接触、支撑、形变和状态转换等物理证据。然而,现有的全模态描述生成器主要建模通用的视听语义,往往忽略瞬态或空间局部的物理证据。我们提出了一个统一的物理感知视听描述生成框架,涵盖数据构建、训练和评估。首先,我们构建了一个数据构建流程,用于识别富含物理信息的片段,并利用 OmniFysics-Agent 协调音频、视觉和物理感知工具,以收集时空对齐且可追溯的跨模态证据;在 Agent 内部,一个在约 200 万图像级样本上微调的物理感知模型(PPM)作为专用工具,用于提取物体交互和状态变化线索。其次,我们构建了 Daily-Physics 50K 数据集,并引入了基于证据的 OmniPhysCap(OPC)基准,以评估从生成的描述中恢复物理和跨模态证据的能力。最后,我们使用由此产生的数据训练 OmniFysics-Captioner。我们的描述生成器在视听描述生成方面与 Gemini 3.1 Pro 相当,在多个视频描述生成基准上取得了最先进的结果,并大幅优于其他开源模型。消融实验表明,PPM 证据改善了物理覆盖,并产生了更细粒度、更可靠的跨模态描述。
英文摘要
Building omni-modal models with physical intelligence requires fine-grained supervision that captures physical evidence such as contact, support, deformation, and state transitions. However, existing omni-modal captioners primarily model general audiovisual semantics and often overlook transient or spatially localized physical evidence. We present a unified framework for physics-aware audiovisual captioning spanning data construction, training, and evaluation. Firstly, we build a data construction pipeline that identifies physics-rich clips and leverages OmniFysics-Agent to coordinate audio, visual, and physical-perception tools for collecting spatiotemporally aligned and traceable cross-modal evidence; within the Agent, a physical perception model (PPM) fine-tuned on approximately 2M image-level samples serves as a dedicated tool for extracting object-interaction and state-change cues. Secondly, we build the Daily-Physics 50K dataset and introduce the evidence-driven OmniPhysCap (OPC) benchmark to evaluate the recovery of physical and cross-modal evidence from generated captions. Finally, we train OmniFysics-Captioner from the resulting data. Our Captioner matches Gemini 3.1 Pro on audiovisual captioning, achieves state-of-the-art results on multiple video-captioning benchmarks, and substantially outperforms other open-source models. Ablations show that PPM evidence improves physical coverage and produces finer-grained, more reliable cross-modal descriptions.
CommentsFysics AI Technical Report