可交互世界的精确编辑与灵活引用
Precise Editing and Flexible Referencing for Interactable Worlds
浏览论文内容
中文总结 AI 辅助
EditWorld通过门控因果注意力和稀疏上下文机制,实现可交互世界中的精确编辑与灵活引用,在WBench-Editing上以73.8总分和80.0编辑分领先。
中文摘要 AI 辅助
我们提出了EditWorld,一个用于可交互世界中精确编辑和灵活引用的视频世界模型。现有的视频世界模型主要关注导航,允许用户探索生成的世界,但对现有世界内容的修改控制有限。EditWorld通过自回归生成过程中流式传输编辑指令和参考图像,将世界建模从探索扩展到精确修改。为支持这些能力,EditWorld引入了门控因果注意力来处理随时间变化的编辑条件和参考图像,并采用稀疏上下文机制来维持有界的历史上下文以进行长时推理。我们进一步采用联合自回归和双向训练,配合退火自重采样,并构建了专门的数据合成和标注流程,为世界编辑提供监督。我们还提出了WBench-Editing来系统评估流式世界编辑能力。EditWorld在WBench-Editing上取得了最佳整体性能,总分为73.8,编辑得分为80.0,在编辑相关指标上大幅超越现有方法。
英文摘要
We present EditWorld, a video world model for precise editing and flexible referencing in interactable worlds. Existing video world models primarily focus on navigation, letting users explore generated worlds but offering limited control over how existing world content is modified. EditWorld extends world modeling from exploration to precise modification by streaming editing instructions and reference images during autoregressive generation. To support these capabilities, EditWorld introduces Gated Causal Attention for temporally varying editing conditions and reference images, together with a Sparse Context mechanism that maintains a bounded historical context for long-horizon inference. We further adopt joint autoregressive and bidirectional training with annealed self-resampling, and construct a dedicated data synthesis and annotation pipeline that provides supervision for world editing. We also present WBench-Editing to systematically evaluate streaming world editing capabilities. EditWorld achieves the best overall performance on WBench-Editing with an overall score of 73.8 and an editing score of 80.0, substantially outperforming existing methods on editing-related metrics. https://github.com/leoisufa/EditWorld
发表机构
- Nanyang Technological University(南洋理工大学)
- StepFun
机构由 AI 辅助整理,请以论文原文为准。