OmniVR:用于恢复退化历史影片的联合视频-音频条件生成模型
OmniVR: Audio-Video Conditional Generation for Archival Footage Restoration
浏览论文内容
中文总结 AI 辅助
提出首个联合音视频生成修复模型OmniVR,通过三项关键设计解决历史影片的音视频共同退化问题,在多类指标上超越现有方法,还发布了相关基准OmniVRBench
中文摘要 AI 辅助
历史影片存在视觉与音频的共同退化问题,包括模糊、噪声、闪烁、嘶嘶声、削波和信号丢失,但现有方法会分别恢复每种模态,导致存在质量差距和跨模态不一致问题。我们提出OmniVR,这是首个联合音视频生成修复模型。OmniVR基于220亿参数的音视频生成主干构建,将修复问题表述为统一多模态DiT内的条件生成:将低质量视频和音频编码为潜在条件,结合固定修复提示,通过联合去噪恢复视觉结构、时间运动和声学细节,所有操作在一个协调目标下完成。该模型的适配包含三项关键设计:(1)联合音视频退化流水线,基于互联网收集的数据模拟真实老影片的特征;(2)架构保留的文本到音视频(T2AV)到音视频到音视频(AV2AV)的过渡,结合提示退火以最大程度保留生成先验;(3)首帧图像到视频(I2V)锚定,搭配损失重加权和波形监督,用于长视频 extrapolation 和音频保真度。我们还提出OmniVRBench,这是首个在200个真实历史片段上评估音视频修复的基准,评估维度包括视觉质量、音频质量、时间一致性和音视频同步性。OmniVR在所有6项视觉指标上均超越所有现有方法,实现最佳音频质量,还能生成自然着色,是首个同时解决这三个方面的方法。代码和权重将公开发布,项目页面:this https URL
英文摘要
Archival footage often suffers from coupled visual and acoustic degradations, yet most restoration systems process the two modalities separately. To address this problem, we present OmniVR, the first systematic framework for joint audio-video restoration, covering data construction, model adaptation, efficient inference, and evaluation. We construct a high-quality audio-video corpus with detailed captions and use a joint degradation pipeline to produce aligned clean and degraded pairs. Using these pairs, we adapt a pretrained text-to-audio-video model (T2AV) by introducing degraded audio-video conditions (TAV2AV), then progressively replace sample captions with a fixed restoration prompt while retaining caption/null rehearsal. The resulting AV2AV model requires no user-provided text. Under a compatible residual-learning model, we prove that this condition-annealing schedule reduces gradient variance and expected restoration risk relative to direct fixed-prompt adaptation at the same training budget. For efficient deployment, OmniVR-Flash combines reduced-resolution video conditioning, MeanFlow-based one-step distillation, and Turbo VAE, achieving approximately 38 fps at 1K and 18 fps at 2K on a single B200 GPU. We further introduce OmniVRBench to evaluate four complementary dimensions: visual quality, audio quality, temporal consistency, and audio-visual synchrony. OmniVR achieves state-of-the-art results on public benchmarks and OmniVRBench. Data, code, and model weights will be released. Project Page: https://xin1u.github.io/OminiVR_PAGE/
发表机构
- University of Science and Technology of China(中国科学技术大学)
- JD Explore Academy(京东探索研究院)
机构由 AI 辅助整理,请以论文原文为准。