发表机构
Kuaishou Technology(快手科技)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出工业规模的LPM用于视频修复,通过统一系统解决UGC退化问题,含数据工程、模型训练和推理。其架构等机制实现高保真修复,已在快手生产应用,节省带宽成本,还能集成产品,证明该方法实用、可扩展且具成本效益。
AI 中文摘要
我们提出了大型处理模型(LPM),这是一个基于扩散的生成框架,用于在复杂的自然退化情况下进行逼真的视频修复。据我们所知,LPM是第一个在工业规模上部署的生成式视频修复模型。LPM通过一个统一的系统来解决用户生成内容(UGC)中的各种退化问题,该系统包括大规模数据工程、基础模型训练和高效推理。其增强的架构、渐进训练策略和时间金字塔推理机制共同实现了对UGC平台上广泛内容分布中任意长视频的高保真、时间一致的修复。LPM已在快手投入生产,模型处理的视频占总观看时间的约45%,在关键体验质量指标上持续改进。除了感知增强,LPM还带来了显著的系统级效益:在可比的感知质量下,相对于快手内部编解码器,它将比特率降低了20%,每年节省数亿带宽成本。其低服务成本还使其能够集成到Kling等产品中,表明生成式修复对于大规模视频处理可以是实用、可扩展且具有成本效益的。
英文摘要
We present the Large Processing Model (LPM), a diffusion-based generative framework for photorealistic video restoration under complex, in-the-wild degradations. To our knowledge, LPM is the first generative video restoration model deployed at industrial scale. LPM addresses the diverse degradations in user-generated content (UGC) through a unified system encompassing large-scale data engineering, foundation-model training, and efficient inference. Its enhanced architecture, progressive training strategy, and temporal-pyramid inference mechanism jointly enable high-fidelity, temporally consistent restoration of arbitrarily long videos across the broad content distribution encountered on UGC platforms. LPM has been deployed in production at Kuaishou, where videos processed by the model account for approximately 45% of total viewing time, delivering consistent improvements across key quality-of-experience metrics. Beyond perceptual enhancement, LPM delivers substantial system-level benefits: at comparable perceptual quality, it reduces bitrate by 20% relative to Kuaishou's in-house codec, yielding annual bandwidth cost savings on the order of hundreds of millions. Its low serving cost also enables integration into products such as Kling, demonstrating that generative restoration can be practical, scalable, and cost-effective for large-scale video processing.
Comments21 pages, 7 figures