多模态大语言模型的离散标记化:综合综述
Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- Jilin University(吉林大学)
- Nanjing University(南京大学)
- University of California at Merced(加州大学默塞德分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该综述首次系统分类并分析面向大语言模型的离散标记化方法,涵盖8种VQ变体,探讨其算法原理、训练动态及在单模态与多模态系统中的集成,指出码本崩溃等挑战并展望动态量化等方向。
AI中文摘要:
大型语言模型(LLM)的快速发展加剧了对有效机制的需求,以将连续的多模态数据转换为适合基于语言处理的离散表示。以向量量化(VQ)为核心方法的离散标记化,既提供了计算效率,又与LLM架构兼容。尽管其重要性日益增长,但目前缺乏一项系统性地审视基于LLM系统中VQ技术的综合综述。本工作通过呈现首个针对LLM设计的离散标记化方法的结构化分类和分析来填补这一空白。我们对8种代表性的VQ变体进行分类,这些变体跨越经典和现代范式,并分析其算法原理、训练动态以及与LLM流水线的集成挑战。在算法层面研究之外,我们从无LLM的经典应用、基于LLM的单模态系统和基于LLM的多模态系统方面讨论现有研究,强调量化策略如何影响对齐、推理和生成性能。此外,我们识别了关键挑战,包括码本崩溃、不稳定的梯度估计和模态特定编码约束。最后,我们讨论了新兴研究方向,如动态和任务自适应量化、统一标记化框架以及受生物学启发的码本学习。本综述弥合了传统向量量化与现代LLM应用之间的差距,为开发高效且可泛化的多模态系统提供了基础参考。持续更新的版本可在以下网址获取:此https URL。
英文摘要:
The rapid advancement of large language models (LLMs) has intensified the need for effective mechanisms to transform continuous multimodal data into discrete representations suitable for language-based processing. Discrete tokenization, with vector quantization (VQ) as a central approach, offers both computational efficiency and compatibility with LLM architectures. Despite its growing importance, there is a lack of a comprehensive survey that systematically examines VQ techniques in the context of LLM-based systems. This work fills this gap by presenting the first structured taxonomy and analysis of discrete tokenization methods designed for LLMs. We categorize 8 representative VQ variants that span classical and modern paradigms and analyze their algorithmic principles, training dynamics, and integration challenges with LLM pipelines. Beyond algorithm-level investigation, we discuss existing research in terms of classical applications without LLMs, LLM-based single-modality systems, and LLM-based multimodal systems, highlighting how quantization strategies influence alignment, reasoning, and generation performance. In addition, we identify key challenges including codebook collapse, unstable gradient estimation, and modality-specific encoding constraints. Finally, we discuss emerging research directions such as dynamic and task-adaptive quantization, unified tokenization frameworks, and biologically inspired codebook learning. This survey bridges the gap between traditional vector quantization and modern LLM applications, serving as a foundational reference for the development of efficient and generalizable multimodal systems. A continuously updated version is available at: https://github.com/jindongli-Ai/LLM-Discrete-Tokenization-Survey.