arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLaMA 3.1 8B中结构感知数值推理的机制可解释性

Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B

Rahul Chowdhury, Timothy A Rupprecht, Senhao Cao, Jiahao Liu, Octavia Camps, David Bau, Pu Zhao, Yanzhi Wang

arXiv 2608.18419首次发表:更新:

发表机构

Northeastern University; EmbodyX Inc.(东北大学; EmbodyX公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究从机制可解释性视角探究LLaMA 3.1 8B,通过构建需捕捉结构的数值序列任务,发现其可无监督计算存储一阶差分,还揭示其通过类诱导回路机制完成数值推理。

AI 中文摘要

近期研究表明,大型语言模型(LLMs)具备强大的数值序列建模能力,在时间序列预测方面展现出应用前景。尽管LLMs表现出上下文学习能力,但其完成时间序列预测的机制仍不明确,具体而言,它们是否真正理解底层结构——至少需要对数字序列的一阶差分进行推理。为研究这一问题,我们从机制可解释性视角探究LLaMA 3.1-8B。机制可解释性是一个新兴领域,关注对LLMs等神经网络所学习算法的逆向工程。为评估LLaMA的数值序列建模能力并推动机制可解释性分析,我们构建了一项无法仅靠表面模式解决的序列建模任务:具体而言,我们采样n个随机数并对其进行偏移重复。研究发现,LLaMA在该任务上表现出色,表明其能够捕捉底层结构。为理解其实现机制,我们开展了探测实验和基于激活补丁的反事实分析。探测实验显示,该模型在无明确监督的情况下计算并存储内部表征中的一阶差分,说明其会追踪序列的结构信息;激活补丁分析揭示,LLaMA通过类似诱导回路的机制检索相关一阶差分,随后将其与当前值相加。值得注意的是,本研究是首批在LLMs中识别出此类概念诱导形式的研究之一。

英文摘要

Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-series prediction. While LLMs display in-context learning capabilities, the mechanisms with which they accomplish time-series prediction remain unclear. Specifically, whether they truly understand the underlying structure, which at a minimum requires reasoning over first differences in the sequence of numbers. To study this, we investigate Llama 3.1-8B from a mechanistic interpretability point of view. Mechanistic interpretability is an emerging field concerned with the reverse engineering of the algorithms learned by neural networks such as LLMs. To assess Llamas' numerical sequence modeling capabilities and to facilitate our mechanistic interpretability analysis, we create a sequence modeling task that cannot be solved without picking up structural cues. Specifically, we sample n random numbers and repeat them with an offset. We find that Llama displays strong performance on our tasks suggesting that it can pick up on the underlying structure. To understand the mechanisms that allow it to do so, we perform probing experiments and activation patching based counterfactual analysis. Probing reveals that the model computes and stores first differences in its internal representations without explicit supervision, indicating that it tracks structural information about the sequence. Activation patching reveals that Llama retrieves the relevant first-difference with a mechanism similar to an induction circuit and subsequently adds it to the current value. Notably, our work represents one of the first studies to identify this form of concept induction in LLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑