arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2505.05422cs.CVcs.AIcs.CL

TokLIP:将视觉令牌与CLIP结合用于多模态理解与生成

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation

  • ARC Lab, Tencent PCG(腾讯PCG广告实验室)
  • City University of Hong Kong(香港城市大学)
  • Zhejiang University(浙江大学)
  • NLPR & MAIS, Institute of Automation, CAS, Beijing(自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

Haokun Lin, Teng Wang, Yixiao Ge, Yuying Ge, Zhichao Lu, Ying Wei, Qingfu Zhang, Zhenan Sun, Ying Shan

更新

AI总结:

针对现有基于令牌的多模态工作训练开销大、理解性能有限的问题,提出TokLIP视觉令牌器,通过解耦理解与生成目标并融入CLIP语义,实现高效数据利用与双任务性能提升,支持自回归训练。

AI中文摘要:

Chameleon和Emu3等开创性的基于令牌的工作为多模态统一奠定了基础,但因缺乏高级语义而面临训练计算开销大、理解性能有限的挑战。本文提出TokLIP,这是一种视觉令牌器,通过将矢量量化(VQ)令牌语义化并融入CLIP级语义来增强理解能力,同时支持使用标准VQ令牌进行端到端多模态自回归训练。TokLIP集成了低级离散VQ令牌器和基于ViT的令牌编码器,以捕捉高级连续语义。与之前离散化高级特征的方法(如VILA-U)不同,TokLIP解耦了理解和生成的训练目标,允许直接应用先进的VQ令牌器,无需定制量化操作。实证结果表明,TokLIP实现了卓越的数据效率,赋予视觉令牌高级语义理解能力,同时增强了低级生成能力,使其非常适合自回归Transformer处理理解和生成任务。代码和模型可在https://github.com/TencentARC/TokLIP获取。

英文摘要:

Pioneering token-based works such as Chameleon and Emu3 have established a foundation for multimodal unification but face challenges of high training computational overhead and limited comprehension performance due to a lack of high-level semantics. In this paper, we introduce TokLIP, a visual tokenizer that enhances comprehension by semanticizing vector-quantized (VQ) tokens and incorporating CLIP-level semantics while enabling end-to-end multimodal autoregressive training with standard VQ tokens. TokLIP integrates a low-level discrete VQ tokenizer with a ViT-based token encoder to capture high-level continuous semantics. Unlike previous approaches (e.g., VILA-U) that discretize high-level features, TokLIP disentangles training objectives for comprehension and generation, allowing the direct application of advanced VQ tokenizers without the need for tailored quantization operations. Our empirical results demonstrate that TokLIP achieves exceptional data efficiency, empowering visual tokens with high-level semantic understanding while enhancing low-level generative capacity, making it well-suited for autoregressive Transformers in both comprehension and generation tasks. The code and models are available at https://github.com/TencentARC/TokLIP.

补充信息

↑