TRaM-VSR:用于一步扩散视频超分辨率的重要性感知令牌路由与合并
TRaM-VSR: Importance-Aware Token Routing and Merging for One-Step Diffusion Video Super-Resolution
- Advanced Micro Devices Inc.(超威半导体公司)
- Computer Vision Lab, CAIDAS & IFI, University of Würzburg(维尔茨堡大学CAIDAS与IFI计算机视觉实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对视频超分辨率中扩散模型计算成本高及现有方法存在问题,提出TRaM-VSR框架,通过融合线索估计令牌重要性,经离线规划器校准调整,在路由组内分处理关键与非关键令牌,实现加速推理且保持质量和一致性。
AI中文摘要:
使用大规模扩散Transformer先验的视频超分辨率(VSR)实现了卓越的感知质量,但由于处理密集时空令牌序列的二次计算成本,通常不实用。现有面向效率的方法存在不可逆转的细节损失和时间闪烁风险,在一步扩散模型中尤为明显。为解决此问题,我们提出了TRaM-VSR,一种用于自适应令牌分配的令牌路由与合并框架,利用上下文感知视频先验和网络级先验。首先,通过融合运动敏感的时间线索与语义文本相似性来估计令牌重要性,分离动态对象和结构边界。接下来,由离线规划器进一步校准和调整此重要性,以指导跨最佳分组网络块的路由。在技术上,在每个路由组内,结构关键令牌在高保真本地流中处理,而信息较少的令牌聚合到紧凑的全局流中,两者均由网络深度调制并与扩散模型的多粒度性质对齐。广泛实验表明,TRaM-VSR在保持最先进重建质量和强大时间一致性的同时,显著加速了推理。代码可在该https URL获取。
英文摘要:
Video super-resolution (VSR) using large-scale Diffusion Transformer (DiT) priors achieves exceptional perceptual quality but is often impractical due to the quadratic computational cost of processing dense spatio-temporal token sequences. Existing efficiency-oriented methods risk irreversible detail loss and temporal flickering, a vulnerability especially pronounced in one-step diffusion models. To address this, we propose TRaM-VSR, a Token Routing and Merging framework for adaptive token allocation, leveraging both context-aware video priors and network-level priors. First, token importance is estimated by fusing motion-sensitive temporal cues with semantic text similarity, isolating dynamic objects and structural boundaries. Next, this importance is further calibrated and adjusted by an offline planner to guide routing across optimally grouped network blocks. Technically, within each routed group, structurally critical tokens are processed in a high-fidelity local stream, while less informative tokens are aggregated into a compact global stream, both modulated by network depth and aligned with the multigranular nature of diffusion models. Extensive experiments show that TRaM-VSR accelerates inference significantly while preserving state-of-the-art reconstruction quality and robust temporal consistency. The code is available at https://github.com/Ree1s/TRaM-VSR.