arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Puffin-World:基于原生3D世界状态扩展统一多模态模型

Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States

Kang Liao, Yihang Luo, Xiao-Ming Wu, Linyi Jin, Size Wu, Chunyu Lin, Yao Zhao, Fei Wang, Wei Li, Chen Change Loy

arXiv 2609.04196首次发表:更新:

发表机构

S-Lab, Nanyang Technological University; University of Michigan; Beijing Jiaotong University; ACE Robotics(南洋理工大学S-Lab; 密歇根大学; 北京交通大学; ACE机器人公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究提出无需外部离线模块的Puffin-World统一多模态模型,联合建模三类原生3D世界状态与Omni-Camera表示,构建含1500万三元组的Puffin-16M数据集,支持多任务协同的闭环应用并发布相关资源。

AI 中文摘要

我们提出Puffin-World,这是一种无需依赖外部离线模块、集成物理理解、空间模拟及3D世界生成与重建的统一多模态架构。为可靠构建并与3D世界交互,该框架联合建模三类原生世界状态:物理(重力场与纬度)、几何(深度)、外观(图像),以及支持多样任务与灵活运动的统一Omni-Camera表示。除建模上述状态外,我们引入跨未来帧传播物理动力学的策略。通过将绝对相机属性锚定至真实世界,Puffin-World实现物理一致且视觉稳定的世界生成。我们进一步在单一生成过程中耦合外观与几何,联合合成每一个未来视图并重建其底层几何。该统一范式支持需多任务协同的交错闭环应用,包括模仿及自校准世界探索。为将Puffin-World扩展至复杂场景,我们构建Puffin-16M,包含1500万组视觉-语言-相机三元组及100万条含多样且具挑战性运动的轨迹。为推动该领域进一步研究,我们发布了代码、模型及数据集。

英文摘要

We propose Puffin-World, a unified multimodal architecture that integrates physical understanding, spatial simulation, and 3D world generation and reconstruction without relying on external offline modules. To reliably construct and interact with 3D worlds, our framework jointly models three native world states: physics (gravity field and latitude), geometry (depth), and appearance (image), together with a unified Omni-Camera representation that supports diverse tasks and flexible motions. Beyond modeling these states, we introduce a strategy for propagating physical dynamics across future frames. By grounding absolute camera properties in the real world, Puffin-World enables physically consistent and visually stable world generation. We further couple appearance and geometry within a single generative process, jointly synthesizing each future view and reconstructing its underlying geometry. This unified paradigm enables interleaved closed-loop applications requiring synergy across multiple tasks, including mimic and self-calibrated world exploration. To scale Puffin-World to complex scenarios, we construct Puffin-16M, comprising 15 million vision-language-camera triplets and 1 million trajectories featuring various and challenging motions. To foster further research in this area, we released the code, models, and datasets.

CommentsProject Page: https://kangliao929.github.io/projects/puffin-world/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑