超越令牌:大型语言与视觉-语言模型的解码方法综述
Beyond Tokens: A Survey on Decoding Methods for Large Language and Vision-Language Models
- Emory University(埃默里大学)
- Illinois Institute of Technology(伊利诺伊理工大学)
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该综述针对LLMs与LVLMs,梳理其解码方法的三种新兴范式,强调解码方法在确保模型输出与用户意图一致上的高效性,同时指出挑战并展望未来方向。
AI中文摘要:
大型语言模型(LLMs)与大型视觉-语言模型(LVLMs)已展现出令人瞩目的生成能力,但确保其输出与用户意图一致仍是一项挑战。尽管多数现有方法在训练阶段解决该问题,但推理阶段方法如解码方法提供了更高效、可扩展的解决方案。解码方法通过引导令牌级选择、执行序列级生成或并行生成令牌以加速过程,来控制模型生成。本综述从近期LLMs与LVLMs解码方法的研究中识别出三种新兴范式,对这些方法进行系统综述,强调当前存在的挑战,并探讨潜在的未来研究方向。本研究旨在凸显解码方法的高效性与有效性,并提供其应用的实用视角。关于LLMs与LVLMs解码方法的论文列表及更多资源可在此处获取:this https URL。
英文摘要:
Large language models (LLMs) and large vision-language models (LVLMs) have demonstrated impressive generative capabilities, yet ensuring their outputs align with user intent is still challenging. While most existing approaches address this issue at the training stage, inference-time approaches like decoding methods offer a more efficient and scalable solution. Decoding methods control model generation by guiding token-level selection, performing sequence-level generation, or generating tokens in parallel to accelerate the process. In this survey, we identify three emerging paradigms from recent works on decoding methods for LLMs and LVLMs, provide a systematic review of these methods, highlight ongoing challenges, and discuss potential future research directions. Our goal is to underscore the efficiency and effectiveness of decoding methods and offer a practical view of their applications. Paper lists and more resources on decoding methods for LLMs and LVLMs can be found at https://github.com/wang2226/Awesome-LLM-Decoding.