arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.07469cs.AI

COMPASS:在语言模型中定位推理所在之处

COMPASS: Finding Where Reasoning Lives in Language Models

Pratyay Dutta, Kowshik Thopalli, Vivek Narayanaswamy

首次发表
浏览论文内容

中文总结 AI 辅助

COMPASS通过利用模型自身直接回答的正确性作为信号,在推理时引导少数注意力头沿正确性方向激活,显著提升数学推理性能并减少生成令牌。

中文摘要 AI 辅助

明确地引出推理能够显著提升大语言模型(LLM)的性能。现有方法需要预先定义推理的特征,无论是通过思维链(CoT)提示设计、对比思维链方向,还是通过稀疏自编码器(SAE)导出的推理特征。对于具有可验证答案的数学推理,我们表明一个更简单的信号就足够了,即模型自身直接回答尝试的正确性。该信号产生一个潜藏方向,能够引出推理。这个方向在大多数注意力头的激活中是可解码的,但只有一小部分头可以被有效干预。我们引入COMPASS,一种推理时引导方法,该方法使用对数空间(logit-space)归因分数识别这些头,并沿着正确性方向引导它们的激活,仅需要每个头的激活统计信息。在三个模型家族和多个数学基准上,COMPASS优于我们比较的激活引导基线,平均将GSM8K准确率提高16个百分点,并以生成令牌减少20-70%的方式接近CoT准确率。干预无需重新拟合即可迁移到未见过的基准,消融实验表明正确性方向和携带该方向的小部分头都是必要的,且效果集中在极少数头上。

英文摘要

Explicitly eliciting reasoning substantially improves LLM performance. Existing approaches require a predefined characterization of reasoning, whether through CoT prompt design, contrastive CoT directions, or via SAE derived reasoning features. For mathematical reasoning with verifiable answers, we show that a much simpler signal suffices, which is the correctness of the model's own direct answer attempts. This signal yields a latent direction that elicits reasoning. This direction is decodable within the activations of most attention heads, but only a small subset of them can be effectively intervened. We introduce COMPASS, an inference-time steering method that identifies these heads using a logit-space attribution score and steers their activations along the correctness direction, requiring only per-head activation statistics. Across three model families and multiple math benchmarks, COMPASS outperforms the activation-steering baselines we compare against, improves GSM8K accuracy by 16 percentage points on average, and approaches CoT accuracy with 20-70\% fewer generated tokens. Interventions transfer without re-fitting to unseen benchmarks, and ablations show that both the correctness direction and the small set of heads carrying it are necessary, with the effect concentrated in remarkably few heads.

发表机构

  • University of California, Riverside(加州大学河滨分校)
  • Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑