arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理错误在大语言模型(LLMs)残差流轨迹中具有特定区域与方向

Reasoning Errors Have a Region and a Direction in the Residual-Stream Trajectory of LLMs

Hamed Damirchi, Ignacio Meza De la Jara, Damith Ranasinghe, Yuhang Liu, Javen Shi

arXiv 2608.05660首次发表:更新:

发表机构

Australian Institute for Machine Learning; Adelaide University; Naval Group Pacific; Responsible AI Research Centre(澳大利亚机器学习研究所; 阿德莱德大学; 太平洋海军集团; 负责任人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型推理错误检测的权衡问题,提出三流检测器,在推理基准上提升选择准确率,还适用于事实类任务,验证了运动、区域、方向信号的互补性。

AI 中文摘要

随着语言模型越来越多地用于需要可验证推理的任务,可靠区分合理推理与有缺陷的推理已成为重要的实际问题。近期基于轨迹的方法在分层残差流位移中寻找该信号,这种位移捕捉了表征如何变化,同时衰减了一些稳定的、标记特定的信息。然而,位移忽略了更新的起源状态,而恢复完整状态则有重新引入易出现捷径信息的风险。我们识别出这种权衡,提出一种三流检测器,将运动与两个受限位置视图相结合:基于向量量化的粗粒度区域读取器,以及在归一化多层状态上的细粒度方向读取器。该设计恢复了足够的状态上下文以解释运动,又无需回到全状态探测。在训练时未见过的推理基准上,我们的方法相比仅用位移的现有技术,选择准确率提升最高达12%;相比单层探测基线,准确率提升21%。尽管仅在推理基准上训练,它也能读取事实补全与事实验证任务,优于所有对比检测器,这表明信号与正确性相关,而非某类推理。消融实验进一步显示,运动、区域和方向提供互补信号。这些结果表明,从状态条件化的运动中读取推理有效性,比单独从静态状态或去语境化轨迹中读取效果更好。

英文摘要

As language models are increasingly used for tasks that require verifiable reasoning, reliably distinguishing sound reasoning from flawed reasoning has become an important practical problem. Recent trajectory-based methods seek this signal in layerwise residual-stream displacements, which capture how representations change while attenuating some stable, token-specific information. However, displacement omits the state from which an update originates, whereas restoring the full state risks reintroducing shortcut-prone information. We identify this trade-off and propose a three-stream detector that combines motion with two restricted views of location. A coarse region reader based on vector quantization and a fine direction reader over normalized multi-layer states. This design restores enough state context to interpret the motion without returning to full-state probing. On reasoning benchmarks unseen during training, our method improves selection accuracy by up to 12% over the displacement-only state of the art and 21% over single-layer probing baselines. Although trained only on reasoning benchmarks, it also reads factual completion and fact verification, ahead of every detector we compare against, which places the signal on correctness rather than on a kind of reasoning. Ablations further show that motion, region, and direction provide complementary signals. These results suggest that reasoning validity is better read from state-conditioned motion than from either static states or decontextualized trajectories alone.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑