arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.08775cs.CLcs.LG

HALO:语言模型的混合自适应潜在推理

HALO: Hybrid Adaptive Latent Reasoning for Language Models

Micah Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究如何用少量自适应计算改进预训练语言模型,提出HALO混合自适应潜在细化方法,结合粗略与选择性第二阶段细化。在基准比较中,HALO成绩最佳,使用细化步骤少,兼具更好的细化分配,计算量也少。

中文摘要 AI 辅助

我们研究如何通过少量自适应额外计算来改进冻结的预训练语言模型。一种简单方法是在主干隐藏状态之上添加额外的细化步骤,但固定的额外细化可能会造成浪费:一步细化头可能太弱,而强制在所有地方进行第二步全序列细化会增加计算量且不提高迁移效果。我们引入了HALO,一种混合自适应潜在细化方法,它将粗略细化阶段与基于令牌评分和单调令牌停止选择的令牌子集上的选择性第二阶段潜在细化相结合。在由MMLU - Pro和GPQA - Diamond构建的主要公共基准比较中,HALO在面向论文的方法中取得了最佳总体平均成绩,优于冻结主干、固定1和固定2。内部分析进一步表明,HALO在使用比固定1更少且比固定2少得多的平均应用细化步骤的情况下,达到了与固定2几乎相同的令牌准确率水平。这些结果表明,关键优势不仅在于更多的细化,还在于更好的细化分配:HALO在获得最强面向论文结果的同时,使用的测量控制器计算量也比任何一个固定基线都少。

英文摘要

We study how to improve a frozen pretrained language model with a small amount of adaptive extra computation. A simple approach is to add additional refinement steps on top of the backbone hidden states, but fixed extra refinement can be wasteful: a one-step refinement head may be too weak, while forcing a second full-sequence refinement step everywhere can increase compute without improving transfer. We introduce HALO, a hybrid adaptive latent-refinement method that combines a coarse refinement stage with selective second-stage latent refinement on a subset of tokens chosen by token scoring and monotonic token halting. On the main public benchmark comparison built from MMLU-Pro and GPQA-Diamond, HALO achieves the best overall average among the paper-facing methods, outperforming the frozen backbone, fixed-1, and fixed-2. Internal analysis further shows that HALO reaches nearly the same token-accuracy level as fixed-2 while using fewer average applied refine steps than fixed-1 and far fewer than fixed-2. These results suggest that the key advantage is not simply more refinement, but a better allocation of refinement: HALO achieves the strongest paper-facing result while also using less measured controller compute than either fixed baseline.

发表机构

  • Lockheed Martin(洛克希德·马丁公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑