发表机构
Kandinsky Lab(坎丁斯基实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出用于多模态生成模型的KVAE分词器家族,含音频、图像、视频专用型号,性能优于多款前沿开源分词器,公开了训练细节与代码。
AI 中文摘要
潜在扩散建模(LDM)是一种主流范式,它利用分词器将输入信号映射为压缩表示。这种依赖关系使分词器成为生成过程的组成部分,因为它会影响学习速度、合成样本的质量,并为后续应用奠定基础。本报告提出了一系列用于音频、图像和视频的KVAE分词器,全部设计用于后续文本条件生成:KVAE-Audio是一种连续全频带48 kHz分词器,具有64个通道、50 Hz的潜在空间;KVAE-3D是两种因果视频分词器,用于4×16×16和4×8×8压缩;KVAE-2D是一种图像模型,将输入压缩8倍,具有32个通道。我们证明,在重建(PSNR、LPIPS、PESQ等)和生成结果的客观指标(Frechet距离、CLIP得分、CLAP得分等)及主观指标(并排评估)上,其性能与Wan-2.2、HunyuanVideo-1.5、FLUX.2、MovieGen、StableAudio和MMAudio等前沿开源VAE分词器相当或更优。考虑到开发难度,我们向社区分享了训练细节、模型选择方法及设计选择的 ablation 实验,代码可通过两个公开URL获取。
英文摘要
Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part of generation process itself, since it affects learning speed, quality of synthesized samples and lay foundation for later applications. This report presents series of KVAE tokenizers for audio, image and video, all designed for subsequent text-conditioned generation: KVAE-Audio, a continuous full-band 48 kHz tokenizer with a 50 Hz latent of 64 channels; KVAE-3D -- two causal video tokenizers for 4x16x16 and 4x8x8 compression; KVAE-2D, an image model, compressing input by factor of 8 with 32 channels. We demonstrate that reconstruction (PSNR, LPIPS, PESQ, etc.) and generation results on objective (Frechet Distance, CLIP score, CLAP score, etc.) and subjective (side-by-side evaluation) metrics matches or surpasses frontier opensource tokenizers, such as VAEs from Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio and MMAudio. Considering difficulty of development, we share with community training details, model selection method and ablation on design choices. The code is publicly available at https://github.com/kandinskylab/kvae and https://github.com/kandinskylab/kvae-audio.