arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

WorldAttention:一种用于交互式视频世界模型的高效注意力架构

WorldAttention: An Efficient Attention Architecture for Interactive Video World Models

Zeyu Zhang, Jinyuan Mao, Dakai An, Wangbo Zhao, Hanfeng Lu, Jiasheng Tang, Yinghao Yu, Wei Wang, Bohan Zhuang

arXiv 2609.34606首次发表:更新:

发表机构

DAMO Academy, Alibaba Group; Zhejiang University; Hong Kong University of Science and Technology; Hupan Lab; TRE, Alibaba Group(阿里巴巴集团达摩院; 浙江大学; 香港科技大学; 湖畔实验室; 阿里巴巴集团TRE)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出WorldAttention,一种结合混合稀疏注意力和分层KV缓存的高效注意力架构,用于交互式视频世界模型,在VBench-Long和InterVBench上超越现有方法。

AI 中文摘要

利用自回归扩散的范式,文本条件下的交互式视频世界模型旨在模拟由文本指令引导的时间上连贯的环境。虽然实现低延迟、长时间生成对于具身智能和基于模拟的规划至关重要,但当前框架主要依赖滑动窗口机制来限制计算复杂度。然而,这种方法固有地牺牲了历史上下文,削弱了长程交互能力。相反,维护完整历史的缓存仍然在计算上不可行且内存密集:注意力的二次复杂度导致过高的计算开销,而KV缓存的线性增长不可避免地导致GPU内存饱和。为了克服这些限制,我们提出了WorldAttention,一种面向系统的注意力架构,通过专用注意力内核和分层KV缓存管理的协同设计实现高效率。首先,我们引入了混合稀疏注意力(HSA),它集成了线性全局注意力并辅以头部自适应稀疏注意力。此外,我们设计了一种分层KV缓存(HKV),它将历史KV对组织成语义索引的页面,跨多层内存存储,实现细粒度检索和受控的GPU驻留。这两种设计由定制内核支持,以有效地将其理论效率转化为实际性能。在VBench-Long和InterVBench上的大量实验表明,WorldAttention持续超越先前的最先进方法,分别在VBench-Long上达到0.9472的主体一致性分数,在InterVBench上达到0.9668。

英文摘要

Leveraging the paradigm of autoregressive diffusion, text-conditioned interactive video world models aim to simulate temporally coherent environments guided by textual instructions. While enabling low-latency, long-duration generation is pivotal for embodied AI and simulation-based planning, current frameworks primarily rely on sliding-window mechanisms to bound computational complexity. However, this approach inherently sacrifices historical context, undermining the long-range interactive capabilities. Conversely, maintaining a full-history cache remains computationally prohibitive and memory-intensive: the quadratic complexity of attention leads to excessive computational overhead, while the linear growth of the KV cache inevitably leads to GPU memory saturation. To overcome these limitations, we propose WorldAttention, a system-oriented attention architecture that achieves high efficiency through the co-design of specialized attention kernels and hierarchical KV cache management. First, we introduce Hybrid Sparse Attention (HSA), which integrates linear global attention supplemented with head-adaptive sparse attention. Additionally, we design a Hierarchical KV Cache (HKV) that organizes historical KV pairs into semantically indexed pages across multi-tier memory, enabling fine-grained retrieval and controlled GPU residency. These two designs are supported by tailored kernels to effectively translate their theoretical efficiency into real-world performance. Extensive experiments on VBench-Long and InterVBench demonstrate that WorldAttention consistently surpasses prior state-of-the-art methods, achieving subject consistency scores of 0.9472 on VBench-Long and 0.9668 on InterVBench, respectively.

CommentsWebsite: https://alibaba-damo-academy.github.io/WorldAttention, Code: https://github.com/alibaba-damo-academy/WorldAttention

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑