arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Sekai2:从世界探索到交互式世界建模

Sekai2: From World Exploration to Interactive World Modeling

Kang He, Wenshuo Peng, Zihui Gao, Jiaming Tan, Kaipeng Zhang, Yongtao Ge

arXiv 2608.09449首次发表:更新:

发表机构

Alaya Lab; Shanghai Innovation Institute; Wuhan University; Tsinghua University(Alaya实验室; 上海创新研究院; 武汉大学; 清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究人员推出多源真实世界视频数据集Sekai2,其具备丰富观测数据与标注,可用于长时程视频生成、相机可控合成及交互式世界模型预训练。

AI 中文摘要

视频世界模型必须捕捉场景如何随时间和不同视点演变,因此,为训练这类模型进行长时程生成和相机控制,需要搭配相机轨迹与时间对齐语义的长视频。现有数据集很少同时具备这三者:大规模网络视频提供了广泛的视觉多样性,但没有轨迹或时间对齐文本;而带姿态标注的数据集通常是短程的或以重建为导向。我们推出Sekai2,这是一个多源真实世界视频数据集,它将Sekai的世界探索素材推向交互式世界建模。该数据集包含来自113个国家或地区的10428个源视频中的128892个片段,总时长2826小时,且特意偏向持续观测:在通用的120秒分解下,有43594个片段达到完整的2分钟,占所有素材的51.4%。每个片段都包含已发布的相机轨迹和分层标注,这些标注将主体运动、环境动态、静态场景内容和相机行为解耦,形成649597个时间锚定片段。关键的是,我们还引入了982个沿非线性轨迹(包含循环和重访)捕获的全景序列,这些重访提供了同一位置在不同时间和视点下的重复观测,为学习持久场景表示、长期空间记忆和几何一致的世界模型提供了必要的监督。数据集规模分析显示,它具备完整的姿态与字幕覆盖、广泛的地理和语义多样性、多样的相机轨迹以及高度非冗余的时间描述。综合这些特性,Sekai2成为可扩展的资源,适用于长时程视频生成、相机可控合成以及交互式世界模型预训练。

英文摘要

Video world models must capture how scenes evolve over time and across viewpoints. Training them for long-horizon generation and camera control therefore benefits from long videos paired with camera trajectories and temporally grounded semantics. Existing corpora rarely offer the three together: large-scale web video provides broad visual diversity but no trajectories or time-aligned text, while pose-annotated datasets are typically short-range or reconstruction-oriented. We introduce Sekai2, a multi-source real-world video dataset that carries the world-exploration footage of Sekai toward interactive world modeling. The release contains 128,892 clips totaling 2,826 hours from 10,428 source videos across 113 countries or regions, and is deliberately weighted toward sustained observation: under a common 120-second decomposition, 43,594 segments reach the full two minutes and account for 51.4% of all footage. Every clip includes a released camera trajectory and hierarchical annotations disentangling subject motion, environment dynamics, static scene content, and camera behavior, resulting in 649,597 temporally grounded segments. Crucially, we further introduce 982 panoramic sequences captured along non-linear trajectories with loops and revisits. These revisits provide repeated observations of the same locations across time and viewpoints, offering essential supervision for learning persistent scene representations, long-term spatial memory, and geometrically consistent world models. Corpus-scale analyses demonstrate complete pose-and-caption coverage, broad geographic and semantic diversity, varied camera trajectories, and highly non-redundant temporal descriptions. Together, these properties make Sekai2 a scalable resource for long-horizon video generation, camera-controllable synthesis, and interactive world-model pre-training.

CommentsSekai2 dataset technical report. Developed at Alaya Lab

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑