发表机构
Stability AI; Karlsruhe Institut für Technologie(Stability AI; 卡尔斯鲁厄理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SemanTok是一种灵活的视频分词器,通过将冻结的DINO特征注入编码器并利用轻量级头从标记前缀重建,实现了高语义对齐与视频保真度,其AR模型在更小规模下即可匹配或超越更大模型。
AI 中文摘要
近期基于视频的世界模型将自回归(AR)预测的可扩展性与扩散模型的视觉质量相结合。场景分词器的选择对于这两者各自的最优性能至关重要,无论是在保真度还是语义方面。灵活长度、从粗到细的分词器恰好提供了这一点:最初的粗粒度标记承载片段的全局语义,而后续标记则进一步细化细节。现有的灵活分词器仅在早期解码器隐藏状态上应用表示对齐(REPA)损失,而这一目标解码器可以部分地从其噪声输入中实现。我们引入了SemanTok,一种灵活的视频分词器,它将冻结的DINO特征馈入其编码器,并添加轻量级头,仅从每个保留的标记前缀重建这些特征。SemanTok在每种AR模型规模下都实现了高语义对齐和视频保真度:一个201M参数的SemanTok AR模型匹配或超越了其3.4倍规模的VideoFlexTok AR模型,而更大的SemanTok AR模型进一步提高了保真度。它在分布外类别上保持了语义对齐,并在每个噪声级别(包括纯噪声)下为解码器提供了更高的语义对齐。它在重建和生成方面均表现良好,其短标记前缀预测成本更低,生成保真度更高,像素细节则推迟到后续标记中处理。
英文摘要
Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fidelity and semantics. Flexible-length, coarse-to-fine tokenizers yield exactly that: the first coarse tokens carry the clip's global semantics while later tokens further specify details. Existing flexible tokenizers only apply a representation-alignment (REPA) loss on early decoder hidden states, a target the decoder can partly meet from its noised input instead. We introduce SemanTok, a flexible video tokenizer that feeds frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each retained token prefix alone. SemanTok achieves high semantic alignment and video fidelity at every AR model size: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model $3.4\times$ its size, and larger SemanTok AR models further improve fidelity. It keeps semantic alignment on out-of-distribution classes and gives the decoder higher semantic alignment at every noise level, including pure noise. It performs well in both reconstruction and generation, and its short token prefixes are cheaper to predict and give better generation fidelity, with pixel detail deferred to later tokens.
Comments29 pages, 22 figures, including references and appendix; 9 pages of main text