发表机构
Alaya Lab(Alaya实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本报告改进AlayaWorld,通过调整条件信号设计等六项修改优化交互式长时世界建模,保留原有架构等核心部分。
AI 中文摘要
本报告介绍了AlayaWorld的改进版本。尽管其骨干架构、分块自回归生成方案和训练数据与之前的版本保持一致,但我们大幅修改了条件信号的表示方式及其向模型的集成方式。新设计遵循一个简单原则:条件信号应在潜在表示和时间结构上与生成内容尽可能匹配。为此,我们进行了两项重大修改:首先,我们将之前基于深度变形的空间内存替换为流式3D点缓存渲染器;其次,我们重新设计了条件流水线,使视觉条件在相同的因果VAE潜在空间中编码,且时间统计与生成视频的时间统计一致。具体而言,新版本引入了六项修改:(1)将静态帧图像条件替换为运动感知潜在条件;(2)将重新渲染的空间内存因果编码为连续序列;(3)在像素空间中对齐时间内存窗口;(4)采用硬内存丢弃,即移除内存标记而非将其归零;(5)统一训练和推理过程中的VAE编码与解码协议;(6)移除相机AdaLN分支,使得视点控制完全通过重新渲染的空间条件提供。
英文摘要
This report presents an improved version of AlayaWorld. While the backbone architecture, chunk-wise autoregressive generation scheme, and training data remain unchanged from the previous release, we substantially revise how conditioning signals are represented and integrated into the model. The new design is guided by a simple principle: conditioning signals should match the generated content as closely as possible in both latent representation and temporal structure. To this end, we make two major changes. First, we replace the previous depth-warping-based spatial memory with a streaming 3D point-cache renderer. Second, we redesign the conditioning pipeline so that visual conditions are encoded in the same causal-VAE latent space, with temporal statistics consistent with those of the generated video. Concretely, the new version introduces six modifications: (1) replacing static-frame image conditioning with motion-aware latent conditioning; (2) causally encoding re-rendered spatial memory as a continuous sequence; (3) aligning the temporal-memory window in pixel space; (4) adopting hard memory dropout that removes memory tokens rather than zeroing them; (5) unifying the VAE encoding and decoding protocol across training and inference; and (6) removing the camera AdaLN branch, such that viewpoint control is provided entirely through the re-rendered spatial condition.
CommentsAuthors are listed alphabetically by the first name and their role. See the contribution section for details