arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文学习放大潜在符号电路

In-Context Learning Amplifies a Latent Symbolic Circuit

Melissa Wessel

arXiv 2609.36265首次发表:更新:

AI 中文总结

本研究揭示大型语言模型中的潜在符号电路在上下文示例积累时被放大,其抽象、归纳、检索三阶段机制在低样本时已可检测,且函数向量可部分替代归纳阶段,依赖完整检索阶段,表明规则遵循能力预先存在于权重中。

AI 中文摘要

大型语言模型能够仅凭少量上下文示例学习抽象规则,但其内部机制如何随示例积累而激活尚不明确。我们追踪了三个模型家族中随样本数量变化的三阶段符号推理电路(抽象、归纳、检索),发现该电路在模型达到高准确率之前就已可检测且功能正常。每个头的因果贡献从1样本到10样本增长达8倍,跨样本激活修补在0样本时将准确率从1%提升至56%,在1样本时从17%提升至88%。在0样本时缩放并注入函数向量可将准确率提升至86%,这很大程度上替代了归纳阶段,但严重依赖于下游检索阶段的完整性。抽象规则遵循的基础设施在没有任何演示之前就已存在于权重中;上下文示例、函数向量及相关干预似乎为同一潜在电路提供输入。

英文摘要

Large language models can learn abstract rules from just a few in-context examples, but how their internal mechanisms activate as examples accumulate is not well understood. We trace a three-stage symbolic reasoning circuit (abstraction, induction, retrieval) across shot counts in three model families and find it is detectable and functional well before the model achieves high accuracy. Per-head causal contribution grows up to 8x from 1- to 10-shot, and cross-shot activation patching raises accuracy from 1% to 56% at 0-shot and 17% to 88% at 1-shot. Function vectors scaled and injected at 0-shot rescue accuracy up to 86%, largely substituting for the induction stage but depending critically on an intact downstream retrieval stage. The infrastructure for abstract rule-following is present in the weights before any demonstrations; in-context examples, function vectors, and related interventions appear to supply input to the same latent circuit.

CommentsAccepted to the Mechanistic Interpretability Workshop at ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑