AI 中文总结
MotionCraft是一种可控视频超分辨率框架,通过结合鲁棒运动融合、潜在世界Transformer与自适应稀疏注意力,实现了时间一致的高质量视频重建,且能灵活权衡时间平滑性与重建保真度。
AI 中文摘要
视频超分辨率(VSR)旨在从低分辨率输入中恢复高保真的高分辨率视频,是移动拍摄、流媒体播放及档案修复等应用的核心技术。现有方法在局部细节保真度、长时空建模、感知真实性与效率之间存在权衡:卷积对齐技术可保留局部结构,但在运动幅度大或退化复杂时表现不佳;基于Transformer的方法能捕捉长程依赖,但需进行架构或算法调整以保持计算可行性;近期基于潜在或扩散的生成器可合成丰富纹理,但需专门的时间约束来维持一致性。本文提出MotionCraft,这是一种可控VSR框架,受世界模型启发将恢复任务建模为感知运动的潜在状态预测,并集成自适应稀疏注意力与明确的用户可访问控制接口。MotionCraft结合了鲁棒运动融合、平衡局部性与针对性非局部交互的潜在世界Transformer,以及紧凑的条件解码器,可在流约束下提供时间一致的高质量重建。实证评估表明,MotionCraft在实现出色重建与感知性能的同时,还能实现时间平滑性与重建保真度之间可预测的权衡。
英文摘要
Video super-resolution (VSR) aims to recover high-fidelity high-resolution videos from low-resolution inputs and is central to applications ranging from mobile capture to streaming and archival restoration. Existing approaches trade off among local-detail fidelity, long-range spatio-temporal modeling, perceptual realism, and efficiency: convolutional alignment techniques preserve local structure but suffer when motion is large or degradations are complex; transformer-based methods capture long-range dependencies yet require architectural or algorithmic adaptations to remain computationally feasible; and recent latent or diffusion-based generators synthesize rich texture but require specialized temporal constraints to maintain coherence. We present MotionCraft, a controllable VSR framework that formulates restoration as motion-aware latent state prediction inspired by world models and integrates adaptive sparse attention with an explicit user-accessible control interface. MotionCraft combines robust motion fusion, a Latent World Transformer that balances locality and targeted non-local interactions, and a compact conditional decoder to deliver temporally consistent, high-quality reconstructions under streaming constraints. Empirical evaluations show that MotionCraft achieves strong reconstruction and perceptual performance while enabling predictable trade-offs between temporal smoothness and reconstruction fidelity.
Comments14 pages, 6 figures