arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LightMTP:轻量级潜在多令牌预测

LightMTP: Lightweight Latent Multi-Token Prediction

Tamara Czinczoll, Julie Kallini, Gerard de Melo, Chen Shani

arXiv 2610.06031首次发表:更新:

发表机构

Hasso Plattner Institute / University of Potsdam; Stanford University; Tel-Aviv University(哈索·普拉特纳研究所 / 波茨坦大学; 斯坦福大学; 特拉维夫大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有MTP方法参数低效或依赖外部模型的问题,提出LightMTP,利用模型自身隐藏状态引导未来令牌表示,仅增加最多1%参数,在通用基准上性能更优,并保持规划、编码和推理的提升。

AI 中文摘要

下一令牌预测(NTP)是大语言模型的标准预训练目标,但它仅为紧邻的下一个令牌提供显式训练信号,这可能导致模型利用局部模式而非捕捉更长距离的结构和思想。多令牌预测(MTP)通过训练模型预测多个未来令牌来解决这一问题。然而,现有的MTP方法通常引入大量新参数,且在下游性能上的提升有限。潜在MTP方法通过将未来令牌编码为向量表示来解决这一效率问题。然而,这些方法通常依赖外部辅助模型进行未来令牌编码。我们提出LightMTP,一种轻量级(即参数高效)的潜在MTP方法,它从模型自身的隐藏状态中引导出未来令牌表示。我们的两种LightMTP变体将监督扩展到更多未来令牌,无需传统MTP的额外计算开销,也无需潜在MTP通常依赖的外部监督。LightMTP最多增加1%的额外参数,在通用语言建模基准上保持更好的性能,并在规划、编码和推理方面取得相似的提升。

英文摘要

Next-token prediction (NTP) is the standard pretraining objective for large language models, yet it provides an explicit training signal only for the immediate next token, which can lead models to exploit local patterns instead of capturing longer-range structure and ideas. Multi-token prediction (MTP) addresses this by training models to predict several future tokens. However, existing MTP methods often introduce a large number of new parameters with limited improvements in downstream performance. Latent MTP approaches address this efficiency issue by encoding future tokens into a vector representation. However, these approaches usually rely on external helper models for future token encoding. We propose LightMTP, a lightweight, i.e., parameter-efficient, latent MTP approach that bootstraps the future token representations from the model's own hidden states. Our two LightMTP variants extend supervision to more future tokens without requiring the additional computational overhead of conventional MTP nor the external supervision latent MTP normally relies on. LightMTP adds at most 1% extra parameters, retains better performance on general language modeling benchmarks, and achieves similar gains in planning, coding, and reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑