发表机构
Pika Labs; POSTECH(皮卡实验室; 浦项科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对实时视频生成中潜视频扩散模型解码慢、内存占用大的问题,提出FlashDecoder,一种基于Transformer的纯视频解码器,通过滚动KV缓存和顺序处理帧实现快速解码,在重建质量匹配卷积解码器的同时大幅提升速度并减少内存占用。
AI 中文摘要
实时视频生成既需要快速去噪,也需要快速解码。然而,当前的潜视频扩散模型依赖于3D卷积解码器,在高分辨率或长视频情况下速度慢且内存占用大。我们引入了FlashDecoder,一种快速、内存高效的纯Transformer视频解码器,逐帧将潜空间解码为像素。在每一步,当前帧通过滚动KV缓存仅关注过去帧的固定大小窗口。固定时间窗口使解码快速且内存可控,与视频长度无关,实现恒定延迟流。由于帧是顺序处理的,无需显式注意力掩码即可强制时间因果关系,能够在高达1080p的分辨率下训练,并与卷积解码器的重建质量相匹配。在Wan2.1和Wan2.2潜空间上,FlashDecoder在重建质量上与每个卷积解码器匹配(例如,在1080p时PSNR为41.55dB对41.49dB),同时在单个H100 GPU上解码速度快3.6倍至4.7倍,内存使用减少多达11倍。通过架构感知推理优化,加速比扩大到了12倍。
英文摘要
Real-time video generation demands fast decoding as much as fast denoising, yet current latent video diffusion models rely on 3D convolutional decoders that are slow and memory-intensive at high resolutions or for long video. We introduce FlashDecoder, a fast, memory-efficient pure-Transformer video decoder that decodes latents to pixels frame by frame. At each step, the current frame attends only to a fixed-size window of past frames through a rolling KV cache. The fixed temporal window keeps decoding fast and memory bounded regardless of video length, enabling constant-latency streaming. Because frames are processed sequentially, temporal causality is enforced without explicit attention masks, enabling training at resolutions up to 1080p and matching the reconstruction quality of convolutional decoders. On the Wan2.1 and Wan2.2 latent spaces, FlashDecoder matches each convolutional decoder in reconstruction quality (e.g., 41.55dB vs. 41.49dB PSNR at 1080p) while decoding 3.6x-4.7x faster with up to 11x less memory on a single H100 GPU. With architecture-aware inference optimizations, the speedup widens to 12x.
CommentsCVPR2026