arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向实时视频超分辨率的流感知扩散模型:基于跨步注意力机制

Streaming-Aware Diffusion for Real-Time Video Super-Resolution via Cross-Step Attention

Harris Partaourides, Sotirios Chatzis

arXiv 2610.11746首次发表:更新:

发表机构

Ethical AI Novelties; Cyprus University of Technology(伦理AI创新公司; 塞浦路斯理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出流感知扩散框架,通过跨步注意力和轨迹耦合扩散调度降低计算复杂度,在REDS4等数据集上实现512×512分辨率下超40 FPS的实时视频超分辨率。

AI 中文摘要

实时视频超分辨率需要在严格延迟约束下实现高时空保真度,这对扩散模型构成挑战,因为扩散模型存在迭代采样成本高、时间协调能力有限的问题。我们提出一种流感知框架,利用视频流的序列结构,将预训练的单图像潜在扩散模型适配为高效的视频超分辨率(VSR)模型。我们的跨步注意力机制可在相邻帧和扩散步骤间复用中间去噪特征,无需显式时间建模即可实现时间信息交换。我们进一步引入轨迹耦合扩散调度,该机制对齐相邻扩散状态,为跨步条件提供更清晰的中间表示,提升时间一致性。这些组件被整合到流推理流水线中,该流水线跨帧增量传播潜在状态,将N帧和S扩散步骤的有效计算复杂度从O(N·S)降低至O(N + S)。在REDS4和YouHQ40-Test数据集上的实验表明,我们的方法在保持帧级稳定性的同时,提升了感知质量和时间真实性,在512×512分辨率下冷启动后帧率超过40 FPS,无需显式时间建模即可实现实时VSR。

英文摘要

Real-time video super-resolution requires high spatio-temporal fidelity under strict latency constraints, challenging diffusion models due to their iterative sampling cost and limited temporal coordination. We propose a streaming-aware framework that adapts pretrained single-image latent diffusion models for efficient video super-resolution (VSR) by exploiting the sequential structure of video streams. Our Cross-Step Attention mechanism reuses intermediate denoising features across adjacent frames and diffusion steps, enabling temporal information exchange without explicit temporal modeling. We further introduce Trajectory-Coupled Diffusion Scheduling, which aligns adjacent diffusion states and provides cleaner intermediate representations for cross-step conditioning, improving temporal coherence. These components are integrated into a streaming inference pipeline that incrementally propagates latent states across frames, reducing the effective computational complexity from $O(N \cdot S)$ to $O(N + S)$ for $N$ frames and $S$ diffusion steps. Experiments on REDS4 and YouHQ40-Test demonstrate improved perceptual quality and temporal realism while maintaining frame-wise stability. Our method achieves over 40 FPS at $512 \times 512$ resolution after cold start, enabling real-time VSR without explicit temporal modeling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑