AI 中文总结
该研究提出潜频有效性(LFV)方法,通过学习VAE特定频谱响应实现快速视频潜空间频谱编辑,其生成的多数算子经评估表现良好,速度约为像素滤波-重新编码的3倍。
AI 中文摘要
在视频VAE潜空间中直接进行频谱编辑可控制噪声、闪烁、平滑度和频率内容,且无需经过解码-滤波-重新编码的流程。然而,视频VAE可能会将像素空间的频带重新分配到潜空间通道中,而潜空间编辑可能会破坏VAE的往返动态。我们提出了潜频有效性(LFV),该方法可学习紧凑的VAE特定频谱响应,仅在提升解码目标保真度且不加剧往返漂移时才应用。LFV遵循从逐频率对角校准器(C1)到全通道混合(CM)的验证选择路径,使跨通道容量成为可控制的逐编辑资源。在覆盖6个频谱族的544个VAE-编辑单元中,LFV生成423个低成本算子:277个由C1处理,146个(占生成算子的34.5%)需要通道混合。在主要的120单元径向扫描中,99/100个生成算子通过了源视频分组的留出评估。在另外5个滤波族中,所有323个生成算子均通过留出评估。完全冻结的OpenVid适配算子(包括验证选择路径系数)在未适配的情况下通过了所有20个测试的CogVideoX和HunyuanVideo生成域单元。所选响应与直接潜空间滤波的延迟相当,速度约为像素滤波-重新编码的3倍。所得映射揭示了不同的VAE机制,包括强通道耦合的CogVideoX响应以及Open-Sora的高频稳定性边界。
英文摘要
Direct spectral editing in video-VAE latents can control noise, flicker, smoothness, and frequency content without a decode--filter--reencode pass. However, video VAEs may redistribute pixel-space frequency bands across latent channels, and latent edits can disrupt VAE round-trip dynamics. We introduce \emph{latent-frequency validity} (LFV), which learns a compact VAE-specific spectral response and deploys it only when it improves decoded-target fidelity without worsening round-trip drift. LFV follows a validation-selected path from a diagonal per-frequency calibrator (C1) to full channel mixing (CM), making cross-channel capacity a controllable per-edit resource. Across 544 VAE--edit cells spanning six spectral families, LFV emits 423 cheap operators: 277 are handled by C1, while 146 (34.5\% of emitted operators) require channel mixing. On the primary 120-cell radial sweep, 99/100 emitted operators pass source-video-grouped held-out evaluation. Across five additional filter families, all 323 emitted operators pass held-out evaluation. Fully frozen OpenVid-fitted operators, including the validation-selected path coefficient, pass all 20 tested CogVideoX and HunyuanVideo generated-domain cells without adaptation. The selected response matches direct latent-filter latency and is about $3\times$ faster than pixel filter--reencode. The resulting maps reveal distinct VAE regimes, including strongly channel-coupled CogVideoX responses and a sharp Open-Sora high-band stability frontier.