发表机构
The Hong Kong Polytechnic University; University of Science and Technology of China(香港理工大学; 中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI生成视频的压缩与编辑需求,提出统一框架,利用冻结生成器作为先验,通过三种时空语义重要性引导技术实现高效压缩与结构保持的提示编辑。
AI 中文摘要
AI生成视频在数量、时长和分辨率上迅速增长,对高效存储和传输提出了日益增长的需求。与从物理世界捕获的自然视频不同,AI生成视频是从学习到的生成分布中采样的样本,其中语义结构对内容一致性至关重要,而许多局部纹理和随机细节可以被合理地重新生成。这一区别表明,压缩应保留语义上重要的时空信息,而不是重建特定生成样本的每个像素。除了重建之外,AI生成视频还产生了对基于提示的编辑的实际需求,用户期望在修改生成内容的同时保留其原始时空语义。受这些观察的启发,我们提出了一个针对AI生成视频的统一压缩与编辑框架,该框架整合了一个冻结的视频生成器作为可复用的生成先验。在此框架内,我们设计了三种时空语义重要性引导的技术,分别解决传输什么、传输多少以及如何使用传输的辅助信息的问题。首先,一种创新选择方法利用时空语义重要性对潜在差异进行投影,使得选中的创新优先考虑语义不变量而非可替换的生成变化。其次,一种帧自适应比特分配方法估计潜在帧的非均匀语义需求,并为需要更强语义保留的帧分配更多创新。第三,一种统一的重建与编辑方法持续调整传输辅助信息的影响,使相同的压缩表示能够为忠实重建提供强有力指导,或作为结构保持的提示驱动编辑的灵活语义锚点。
英文摘要
AI-generated videos are rapidly increasing in volume, duration, and resolution, creating growing demands for efficient storage and transmission. Unlike natural videos captured from the physical world, AI-generated videos are samples from a learned generative distribution, where semantic structures are critical to content consistency, while many local textures and stochastic details can be plausibly regenerated. This distinction suggests that compression should preserve semantically important spatiotemporal information rather than reconstruct every pixel of a particular generative sample. Beyond reconstruction, AI-generated videos also create a practical need for prompt-based editing, where users expect to modify generated content while preserving its original spatiotemporal semantics. Motivated by these observations, we propose a unified compression and editing framework for AI-generated videos that incorporates a frozen video generator as a reusable generative prior. Within this framework, we design three spatiotemporal semantic importance-guided techniques that respectively address what to transmit, how much to transmit, and how to use the transmitted side information. First, an innovation selection method projects the latent discrepancy using spatiotemporal semantic importance, so that the selected innovations prioritize semantic invariants over replaceable generative variations. Second, a frame-adaptive bit allocation method estimates the nonuniform semantic demands of latent frames and allocates more innovations to frames requiring stronger semantic preservation. Third, a unified reconstruction and editing method continuously adjusts the influence of the transmitted side information, enabling the same compressed representation to provide strong guidance for faithful reconstruction or serve as a flexible semantic anchor for structure-preserving prompt-driven editing.