arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

HLA-WM:用于长时程视频世界模型的混合线性注意力

HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models

Zhuokun Chen, Feng Chen, Xi Lin, Xiyu Wu, Jiahao He, Jianfei Cai, Bohan Zhuang

arXiv 2610.05739首次发表:更新:

发表机构

Monash University; The University of Adelaide; Zhejiang University(莫纳什大学; 阿德莱德大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出HLA-WM,一种无需训练的混合线性注意力框架,结合几何引导检索与循环状态计算,解决长时程视频世界模型中GDN的远距离遗忘问题,提升场景一致性并降低内存开销。

AI 中文摘要

长时程视频世界模型需要持久记忆以在长时间生成过程中保持场景一致性。Softmax注意力通过不断增长的KV缓存保留完整的生成历史,而循环线性注意力则将历史压缩为固定大小的状态,内存成本大幅降低。然而,我们识别出Gated DeltaNet(GDN)存在严重的远距离遗忘问题,其中来自遥远但相关场景的信息会被后续状态更新逐渐衰减。为解决这一局限,我们提出HLA-WM,一种无需训练的混合线性注意力框架,结合了粗粒度几何引导检索与细粒度循环线性状态计算。HLA-WM利用GDN的仿射结构缓存紧凑的分块转移摘要,使用相机几何检索与场景相关的历史分块,并将其重组为查询特定的循环状态。在60秒的SANA-WM-Bench上,HLA-WM在不额外训练的情况下,提升了基础自回归生成器的全部六项聚合重访一致性和相机控制指标,包括0.74 dB的PSNR增益和28.5%的旋转误差降低。这些改进在下游细化后依然保持,并泛化到MBench-A,在547个样本上,HLA-WM在所有四个子集和所有评估推理模式下,一致提升了全部三项重访一致性指标。在60秒上下文下,HLA-WM相对于完整KV缓存将历史状态内存减少了12倍,同时推理吞吐量最多降低1.6%。这些结果表明,可选择性寻址的循环记忆能够改善远距离场景回忆,同时保持GDN的效率优势。项目页面:此https URL

英文摘要

Long-horizon video world models require persistent memory to preserve scene consistency over extended rollouts. Softmax attention retains the full generation history through a growing KV cache, whereas recurrent linear attention compresses history into fixed-size states with substantially lower memory cost. However, we identify severe long-range forgetting in Gated DeltaNet (GDN), where information from distant but relevant scenes is progressively attenuated by subsequent state updates. To address this limitation, we propose HLA-WM, a training-free hybrid linear-attention framework that combines coarse-grained geometry-guided retrieval with fine-grained recurrent linear-state computation. HLA-WM exploits the affine structure of GDN to cache compact chunk-wise transition summaries, retrieve scene-relevant historical chunks using camera geometry, and recompose them into query-specific recurrent states. On the $60$-second SANA-WM-Bench, HLA-WM improves all six aggregate revisit-consistency and camera-control metrics of the base autoregressive generator without additional training, including a $0.74$ dB PSNR gain and a $28.5\%$ reduction in rotation error. The improvements persist after downstream refinement and generalize to MBench-A, where HLA-WM consistently improves all three revisit-consistency metrics across all four subsets and all evaluated inference modes over $547$ samples. At a $60$-second context, HLA-WM reduces historical-state memory by $12\times$ relative to full KV caching while incurring at most a $1.6\%$ reduction in inference throughput. These results demonstrate that selectively addressable recurrent memory can improve long-range scene recall while preserving the efficiency advantages of GDN. Project page: https://caesarhhh.github.io/hla-wm/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑