arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24293cs.CV

保留还是丢弃?用于紧凑视频表示的自适应分词器

Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation

Yeonkyeong Lee, Hyunsung Go, Jongmin Kim, Sewoong Lim, Donghoon Lee

首次发表
浏览论文内容

中文总结 AI 辅助

针对传统VAE固定压缩率无法适配视频时空内容复杂度的问题,提出基于Transformer的自适应令牌分词器KATok,结合两种位置预测策略缓解空间错位,在先进压缩率下实现出色的视频重构与生成质量。

中文摘要 AI 辅助

潜在扩散模型已成为高保真图像与视频合成的主流框架,其在紧凑潜在空间中结合变分自编码器(VAE)运行,可在不损害视觉质量的前提下提升计算效率。然而,传统VAE对视频数据并非最优,因为它们采用固定压缩率,无法适应时空内容的变化复杂度。我们提出KATok(Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation),一种基于Transformer的VAE,它包含与潜在令牌联合学习的自适应令牌选择器。通过评估每个令牌的内容丰富度作为保留或丢弃概率,令牌选择器可有效丢弃无信息令牌,自然实现依赖数据的压缩。将自适应分词应用于扩散模型可能引发空间错位,因为令牌丢弃会破坏原始时空结构。为缓解该问题,我们提出两种位置预测策略:级联生成与联合生成,以确保空间一致性。我们的实验表明,该模型在达到先进压缩率的同时,实现了出色的重构与生成质量。对视频数据的进一步分析显示,这种改进主要通过减少时空冗余并移除无信息令牌实现,定量与定性结果均支持这一点。

英文摘要

Latent diffusion models have emerged as a dominant framework for high-fidelity image and video synthesis, operating in compact latent spaces with variational autoencoders (VAEs) to enhance computational efficiency without compromising visual quality. However, conventional VAEs are suboptimal for video data as they employ fixed compression ratios that cannot adapt to the varying complexity of spatio-temporal content. We present KATok (Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation), a transformer-based VAE that incorporates an adaptive token selector which is jointly learned with latent tokens. By evaluating each token's content-richness as keep-or-drop probability, the token selector effectively discards uninformative tokens, naturally allowing data-dependent compression. Applying adaptive tokenization to diffusion models may cause spatial misalignment, as token dropping can disturb the original spatio-temporal structure. To alleviate this issue, we propose two position-prediction strategies: cascaded and joint generation, to ensure spatial consistency. We empirically show that our model achieves strong reconstruction and generation quality at a state-of-the-art compression ratio. Further analysis on video data reveals that this improvement is primarily achieved by reducing spatio-temporal redundancy and removing uninformative tokens, as supported by both quantitative and qualitative results.

发表机构

  • Kakao Corp.(Kakao公司)

机构由 AI 辅助整理,请以论文原文为准。

↑