arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.10519cs.CV

SparSTAR:用于时空自回归视频合成的稀疏注意力机制

SparSTAR: Sparse Attention for SpaceTime AutoRegressive Video Synthesis

Jongbeom Lee, Hyunwoo Yu, Jincheol Yang, Jaemin Choi, Suk-Ju Kang

首次发表
浏览论文内容

中文总结 AI 辅助

针对InfinityStar视频合成中高成本注意力与不可靠稀疏模式问题,提出无需训练的SparSTAR块稀疏注意力,在720p生成任务中实现约1.6倍加速且保持高保真度。

中文摘要 AI 辅助

InfinityStar通过图像和片段金字塔序列将视觉自回归生成扩展到视频领域。然而其不断变化的尺度和跨片段上下文,使得后期尺度的注意力计算成本高昂,且从扩散模型或图像VAR模型复用的稀疏模式不可靠。我们提出SparSTAR,一种针对该场景量身定制的无需训练的块稀疏注意力方法。在每个高成本尺度和注意力头处,SparSTAR对当前查询和键激活的连续键块进行评分,保留所需的条件上下文,并通过仅前向的稀疏路径执行所选块。我们分析了片段内的跨尺度一致性、跨片段边界的模式持久性,以及随着复用跨度涉及越来越远的尺度时的质量下降情况。在所有这些分析中,重要的键块会发生变化,这表明在每个目标尺度重新计算块选择比复用迁移的掩码更可靠。在720p文本到视频和图像到视频生成任务中,SparSTAR保留了所有令牌和细化尺度,同时提供约1.6倍的端到端加速,并保持与密集型InfinityStar接近的VBench和配对输出重建保真度。

英文摘要

InfinityStar extends visual autoregressive generation to video through a sequence of image and clip pyramids. Its changing scale and cross-clip context, however, leave late-scale attention costly and make sparse patterns reused from diffusion or image VAR models unreliable. We introduce SparSTAR, a training-free block-sparse attention method tailored to this setting. At each expensive scale and attention head, SparSTAR scores contiguous key blocks from the current query and key activations, retains required conditioning context, and executes the selected blocks through a forward-only sparse path. We analyze cross-scale consistency within a clip, pattern persistence across clip boundaries, and quality degradation as reuse spans increasingly distant scales. Across these analyses, important key blocks shift, showing that recomputing block selection at each target scale is more reliable than reusing a transferred mask. On 720p text-to-video and image-to-video generation, SparSTAR preserves every token and refinement scale while providing about a 1.6x end-to-end speedup and maintaining VBench and paired-output reconstruction fidelity close to dense InfinityStar.

补充信息

↑