发表机构
Insta360 Research; Institute of Automation Chinese Academy of Sciences; Tsinghua University; Wuhan University; UC Merced(影石创新科技研究院; 中国科学院自动化研究所; 清华大学; 武汉大学; 加州大学默塞德分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究旨在解决全景世界模型的远程记忆挑战,提出PanoWorld,利用旋转等变特性简化相机轨迹,通过DPRC和GMA进行建模与记忆增强,经三阶段训练优化,构建World360数据集验证其有效性,优于其他方法且将公开模型、代码和数据集。
AI 中文摘要
在这项工作中,我们旨在通过利用全向表示的旋转等变特性来应对全景世界模型中的远程记忆挑战,其中旋转可视为隐式几何变换。基于此,我们提出了PanoWorld,它通过固定航向将相机轨迹简化为平移,用于当前动作建模和远程记忆,借助密集全景光线条件(DPRC)和几何感知记忆增强(GMA)。然后,引入了一个三阶段训练管道来逐步优化每个组件。为了在现有数据集相对稳定的大规模空间变化和多样光照条件下更好地评估物理一致性,我们构建了World360数据集,它由通过全景无人机收集的真实世界视频片段和高质量模拟片段组成。在World360上的实验证明了PanoWorld的有效性,其性能大幅优于替代方法。模型、训练代码和数据集将公开可用。更多信息可在项目页面查看。
英文摘要
In this work, we aim to address the challenge of long-range memory in panoramic world models by exploiting the rotation-equivariant property of omnidirectional representations, where rotation can be treated as an implicit geometric transformation.Building on this insight, we propose PanoWorld, which simplifies camera trajectories into translations via fixed headings for both current-action modeling and long-range memory through Dense Panoramic Ray-Conditioning (DPRC) and Geometry-aware Memory Augmentation (GMA).Then, a three-stage training pipeline is introduced to progressively optimize each component. To better evaluate physical consistency under large-scale spatial variations and diverse illumination conditions, where existing datasets are relatively stable, we construct World360, a large-scale dataset consisting of both real-world video clips collected via panoramic unmanned aerial vehicles and high-quality simulated clips generated by AirSim360.Extensive experiments on World360 demonstrate the effectiveness of PanoWorld, outperforming alternative methods by a large margin.Our models, training code, and dataset will be publicly available. More information can be found on our project page: https://lihaoy-ux.github.io/panoworld-page/.
CommentsProject page: https://lihaoy-ux.github.io/panoworld-page/ Code:https://github.com/Insta360-Research-Team/PanoWorld