EchoWM:开放且可进入的全模态世界模型
EchoWM: Open and Enterable Omnimodal World Models
另 6 家 · 查看机构详情
- HKUST(香港科技大学)
- PKU(北京大学)
- Joy Future Academy, JD(京东探索研究院)
- HKU(香港大学)
- THU(清华大学)
- USTC(中国科学技术大学)
- FDU(复旦大学)
- Beihang University(北京航空航天大学)
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究提出EchoWM全模态世界模型,围绕相机意图组织交互,构建互补数据引擎并采用渐进式训练等方法,在公开基准上实现强轨迹跟随与高视觉质量,支持跨视角交互及长序列多模态同步生成。
中文摘要 AI 辅助
我们提出EchoWM,这是一种可进入生成媒体的全模态世界模型,可响应连续导航,同时生成720p视频、环境音、音乐和语音。我们围绕相机意图组织交互:在第一人称场景中,相机意图指定观察者运动;在第三人称场景中,相机-角色动力学从数据中学习,无需特定视图控制器。离散命令与连续位姿被映射到共享的度量尺度相对6自由度轨迹,通过数据集级校准保留异构数据间的运动幅度。为联合学习视听生成与轨迹控制,我们构建了互补数据引擎,并采用渐进式训练后接自回归后训练以实现长序列生成。大量评估显示,EchoWM在公开世界模型基准上实现了强轨迹跟随能力与高视觉质量,支持第一、第三人称跨不同主体交互,并在长序列生成中保持环境音与语音的同步性。
英文摘要
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.