arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MotionCraft:用于视频超分辨率的基于稀疏注意力的潜在世界建模

MotionCraft: Latent World Modeling with Sparse Attention for Visual Upscaling

Rong Fu, Chunlei Meng, Yangchen Zeng, Xiaowen Ma, Yongtai Liu, Wangyu Wu, Shuo Yin, Zijian Zhang, Sicheng Li, Yingrui Ji, Chenhao Wang, Simon Fong

arXiv 2608.08553首次发表:更新:

AI 中文总结

MotionCraft是一种可控视频超分辨率框架,通过结合鲁棒运动融合、潜在世界Transformer与自适应稀疏注意力,实现了时间一致的高质量视频重建,且能灵活权衡时间平滑性与重建保真度。

AI 中文摘要

视频超分辨率(VSR)旨在从低分辨率输入中恢复高保真的高分辨率视频,是移动拍摄、流媒体播放及档案修复等应用的核心技术。现有方法在局部细节保真度、长时空建模、感知真实性与效率之间存在权衡:卷积对齐技术可保留局部结构,但在运动幅度大或退化复杂时表现不佳;基于Transformer的方法能捕捉长程依赖,但需进行架构或算法调整以保持计算可行性;近期基于潜在或扩散的生成器可合成丰富纹理,但需专门的时间约束来维持一致性。本文提出MotionCraft,这是一种可控VSR框架,受世界模型启发将恢复任务建模为感知运动的潜在状态预测,并集成自适应稀疏注意力与明确的用户可访问控制接口。MotionCraft结合了鲁棒运动融合、平衡局部性与针对性非局部交互的潜在世界Transformer,以及紧凑的条件解码器,可在流约束下提供时间一致的高质量重建。实证评估表明,MotionCraft在实现出色重建与感知性能的同时,还能实现时间平滑性与重建保真度之间可预测的权衡。

英文摘要

Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence. We present MotionCraft, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface. MotionCraft combines robust motion fusion, a Latent World Transformer that balances locality and targeted non-local interactions, and a compact conditional decoder to deliver temporally consistent, high-quality reconstructions under streaming constraints. Empirical evaluations show that MotionCraft achieves strong reconstruction and perceptual performance while enabling predictable trade-offs between temporal smoothness and reconstruction fidelity.

Comments14 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑