arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniFysics-Captioner 技术报告:在物理世界中锚定全模态理解以提升描述生成

OmniFysics-Captioner Technical Report: Grounding Omni-Modal Understanding in the Physical World for Better Captioning

Kaixiang Qiu, Minghao Han, Keliang Liu, Yizhou Liu, Jinghan Han, Yue Jiang, Xuecheng Wu, Shunli Wang, Lihua Zhang, Dingkang Yang

arXiv 2609.31714首次发表:更新:

AI 中文总结

针对现有全模态描述生成器忽视物理证据的问题,提出统一框架,通过数据构建、训练和评估,利用物理感知模型提取线索,训练出在多个基准上达到最先进水平的描述生成器。

AI 中文摘要

构建具有物理智能的全模态模型需要细粒度的监督信号,以捕获接触、支撑、形变和状态转换等物理证据。然而,现有的全模态描述生成器主要建模通用的视听语义,往往忽略瞬态或空间局部的物理证据。我们提出了一个统一的物理感知视听描述生成框架,涵盖数据构建、训练和评估。首先,我们构建了一个数据构建流程,用于识别富含物理信息的片段,并利用 OmniFysics-Agent 协调音频、视觉和物理感知工具,以收集时空对齐且可追溯的跨模态证据;在 Agent 内部,一个在约 200 万图像级样本上微调的物理感知模型(PPM)作为专用工具,用于提取物体交互和状态变化线索。其次,我们构建了 Daily-Physics 50K 数据集,并引入了基于证据的 OmniPhysCap(OPC)基准,以评估从生成的描述中恢复物理和跨模态证据的能力。最后,我们使用由此产生的数据训练 OmniFysics-Captioner。我们的描述生成器在视听描述生成方面与 Gemini 3.1 Pro 相当,在多个视频描述生成基准上取得了最先进的结果,并大幅优于其他开源模型。消融实验表明,PPM 证据改善了物理覆盖,并产生了更细粒度、更可靠的跨模态描述。

英文摘要

Building omni-modal models with physical intelligence requires fine-grained supervision that captures physical evidence such as contact, support, deformation, and state transitions. However, existing omni-modal captioners primarily model general audiovisual semantics and often overlook transient or spatially localized physical evidence. We present a unified framework for physics-aware audiovisual captioning spanning data construction, training, and evaluation. Firstly, we build a data construction pipeline that identifies physics-rich clips and leverages OmniFysics-Agent to coordinate audio, visual, and physical-perception tools for collecting spatiotemporally aligned and traceable cross-modal evidence; within the Agent, a physical perception model (PPM) fine-tuned on approximately 2M image-level samples serves as a dedicated tool for extracting object-interaction and state-change cues. Secondly, we build the Daily-Physics 50K dataset and introduce the evidence-driven OmniPhysCap (OPC) benchmark to evaluate the recovery of physical and cross-modal evidence from generated captions. Finally, we train OmniFysics-Captioner from the resulting data. Our Captioner matches Gemini 3.1 Pro on audiovisual captioning, achieves state-of-the-art results on multiple video-captioning benchmarks, and substantially outperforms other open-source models. Ablations show that PPM evidence improves physical coverage and produces finer-grained, more reliable cross-modal descriptions.

CommentsFysics AI Technical Report

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑