发表机构
University of Texas at Austin; Microsoft Research(得克萨斯大学奥斯汀分校; 微软研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对语言模型下一个token预测的误差累积问题,提出分层隐式预测方法,在编码、多步推理基准及推测解码上验证了其有效性。
AI 中文摘要
标准的下一个token预测(NTP)是语言模型预训练的基础,但其教师强制训练范式可能并非长程推理与规划的最优方案。近期如多token预测(MTP)和下一个隐式预测(NextLat)等工作,尝试通过预测多个未来token及隐式空间内的自监督预测来缓解该问题,但这些辅助目标要么视野有限,要么存在多步展开带来的误差累积问题。本文提出分层隐式预测(HiLP),引入辅助的高层抽象隐式变量以减少隐式空间展开中的误差累积效应。实验表明,HiLP可生成更长程的连贯信念状态表示,在编码与多步推理基准上验证了方法的有效性,还能提升推测解码的效率。
英文摘要
While standard Next-Token Prediction (NTP) lays the foundation of language model pre- training, its teacher-forced training paradigm may not be optimal for long-horizon reasoning and planning. Recent works such as Multi-Token Prediction (MTP) and Next-Latent prediction (NextLat) try to mitigate the problem through predicting multiple future tokens and self-supervised prediction in the latent space. However, those auxiliary objectives either have a limited horizon or suffer from compounding error from multi-step rollout. We introduce Hierarchical Latent Prediction (HiLP), which introduces an auxiliary higher-level abstract latent to help reduce the error accumulation effect in latent-space rollouts. Experiments show that HiLP can lead to longer-horizon coherent belief state representation and demonstrate the effectiveness of our method across coding and multi-step reasoning benchmarks, and offers more speculative decoding efficiency.