arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GVCCTurbo:基于码本驱动生成压缩的码率-计算量质量调度

GVCCTurbo: Rate-Compute Quality Scheduling for Codebook Driven Generative Compression

Ziyue Zeng, Dingjie Peng, Xun Su, Hiroshi Watanabe

arXiv 2608.03517首次发表:更新:

AI 中文总结

GVCCTurbo是一种BPP驱动的调度器,分离生成压缩的先验刷新与码本校正,降低解码时间,在低码率下提升压缩性能且兼容多种生成压缩模型。

AI 中文摘要

码本驱动生成压缩采用预训练的图像或视频生成器作为零样本视觉先验,在超低码率下传输紧凑的码本索引以指导重建。当前编解码器将每个有限率校正与一次新的先验评估绑定,因此缩短采样器会移除携带目标相关信息的校正时隙。我们提出GVCCTurbo,一种基于BPP(每像素比特数)的调度器,其将开销较大的先验刷新与码本校正分离:在每个协议上校准一次原子计数工作点和跳隙比后,它将目标码本有效载荷比特率映射为轨迹长度和刷新周期,使BPP成为调度输入而非采样器长度的固定结果。相同的端点预测和有限率控制接口覆盖GVCC风格的整流流视频压缩与DDCM风格的扩散图像压缩,保留零训练部署特性及与未来蒸馏先验的兼容性。原生1080p曲线将完整零样本编解码器置于超低码率 regime 中。在受控720p Wan-GVCC研究中,该调度器将先验评估次数从20次降至9次,实现整个调度系列约44%的实测解码时间减少,仅在高运动内容上产生小幅共享LPIPS(学习感知图像块相似度)代价;在该系列内,均匀刷新稀疏化(纯跳空)是一个边界点,而BPP感知的内部点以少2.9%的码本有效载荷比特数换取在相当LPIPS下始终更高的PSNR(峰值信噪比)。这些结果支持BPP到计算量的调度作为采样器长度调优的可控扩展,无需使分配的工作点主导每个边界点。

英文摘要

Codebook-driven generative compression uses a pretrained image or video generator as a zero-shot visual prior and transmits compact codebook indices to guide reconstruction at ultra-low bitrate. Current codecs tie each finite-rate correction to a fresh prior evaluation, so shortening the sampler also removes correction slots that carry target-dependent information. We propose GVCCTurbo, a BPP-driven scheduler that separates expensive prior refreshes from codebook corrections: after calibrating an atom-count operating point and skip-gap ratio once per protocol, it maps a target codebook-payload bitrate to a trajectory length and refresh period, making BPP a schedule input instead of a fixed consequence of sampler length. The same endpoint-prediction and finite-rate steering interface covers GVCC-style rectified-flow video and DDCM-style diffusion image compression, preserving zero-training deployment and compatibility with future distilled priors. Native 1080p curves position the complete zero-shot codec in the ultra-low-bitrate regime. In a controlled 720p Wan-GVCC study, the scheduler cuts prior evaluations from 20 to 9 for a $\sim\!44\%$ measured decoding-time reduction shared across the whole schedule family, at a small shared LPIPS cost on high-motion content; within that family, uniform refresh thinning (pure-skip) is a boundary point, and the BPP-aware interior point trades $2.9\%$ fewer codebook-payload bits for consistently higher PSNR at comparable LPIPS. These results support BPP-to-compute scheduling as a controllable extension of sampler-length tuning, without requiring the allocated point to dominate every boundary point.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑