发表机构
MoE Key Lab of BIPC; USTC; Shanghai Innovation Institute; Alaya Lab(BIPC教育部重点实验室; 中国科学技术大学; 上海创新研究院; Alaya实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Alaya-EVOKE将持久世界状态外部化并重新设计教师模型,解决交互式世界模型的冲突需求,在WBench等基准上实现最优性能,支持开放式长时序生成。
AI 中文摘要
交互式世界模型必须支持持久记忆、响应式交互和长时序生成,但这些要求对模型提出了相互冲突的需求:在去噪器上下文或键值缓存中维护历史会产生不断增长的成本,迫使在会话长度和保留的记忆之间进行权衡;而低延迟交互依赖于少步生成,其能力受限于其教师模型。EVOKE通过将持久世界状态外部化并重新设计教师模型以支持长时序交互式生成,解决了这两个限制。场景几何结构被维护在一个外部的、按相机索引的世界状态库中,仅从中检索与视图相关的信息,使去噪器上下文在会话增长时保持受限。我们未将教师模型视为固定生成器,而是为其设计了长时序监督:其稀疏注意力结合了分块分组、选定远帧的检索以及线性注意力全局状态,使内存和计算呈线性增长,同时实现了对长时序的监督。这种监督揭示了在短窗口内局部看似合理的内容漂移,而每个分块的条件设置则支持整个序列中的提示变化和事件控制。在自强制回滚下应用的30秒分布匹配目标,将两种能力转移到不使用无分类器引导的三步学生模型,在保持响应式条件设置的同时提高了对长期漂移的抵抗力。凭借受限上下文和循环外部记忆,EVOKE支持开放式、不断演变的生成;在单个H200上,以384×640分辨率,每1.5秒的分块生成耗时2.11秒。作为三步世界模型,EVOKE在WBench上实现了最先进的性能,同时在VBench-Long和VBench-2.0上保持竞争力。
英文摘要
Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher. Alaya-EVOKE (Evoke) addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation. Scene geometry is maintained in an external, camera-indexed world state bank, from which only view-relevant information is retrieved, keeping the denoiser context bounded as the session grows. Rather than treating the teacher as a fixed generator, we design it for long-horizon supervision: its sparse attention combines chunk-wise grouping, retrieval of selected distant frames, and a linear-attention global state, yielding linear growth in memory and compute while enabling supervision over long horizons. Such supervision exposes content drift that stays locally plausible within short windows, while per-chunk conditioning enables prompt changes and event control throughout the sequence. A 30-second distribution-matching objective, applied under self-forced rollouts, transfers both capabilities to a three-step student that uses no classifier-free guidance, improving resistance to long-term drift while preserving responsive conditioning. With bounded context and recurrent external memory, Evoke supports open-ended, continuously evolving generation; on a single H200 at $384\times 640$, each $1.5\,\mathrm{s}$ chunk is generated in $2.11\,\mathrm{s}$. As a three-step world model, Evoke achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.