arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一个共享子电路使大语言模型能够跨任务进行倒计时

A Shared Subcircuit Lets LLMs Count Down Across Tasks

Jacob Dunefsky, Wes Gurnee, Emmanuel Ameisen

arXiv 2607.12279首次发表:更新:

发表机构

Yale University; Anthropic(耶鲁大学; Anthropic)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现Llama-3.1-70B-Instruct中有个“倒计时子电路”可跨任务执行特定任务,先在受控设置中分离它,再研究其表示几何结构,发现与另一模型有相同模式,还通过无监督探测找到其用于多种任务的情况,有助于理解行为推广。

AI 中文摘要

写一个恰好十二个单词的句子、在正确密码子处结束DNA序列、格式化ASCII表格等任务,都需要语言模型跟踪距目标还剩多少令牌。在这项工作中,我们在Llama-3.1-70B-Instruct中识别出执行这些任务的通用机制:一个“倒计时子电路”,它将当前位置与目标长度进行比较并估计剩余时间。我们首先在受控设置中分离出倒计时子电路,然后研究其使用的表示几何结构,发现该子电路使用了在另一个前沿大语言模型的单独任务中识别出的相同模式,最后通过无监督探测找到该子电路用于多种其他任务的情况。我们的工作表明,对子电路进行逆向工程能让我们理解行为如何从单个示例推广到许多不同任务甚至模型。

英文摘要

Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table. These are all tasks that language models can do that requires tracking how many tokens remain before a target. In this work, we identify in Llama-3.1-70B-Instruct a general mechanism for performing these tasks: a "countdown subcircuit" that compares the current position to a goal length and estimates the time remaining until then. We first isolate a countdown subcircuit in a controlled setting, in which the model is tasked with writing a fixed-length sentence ending in a specified word. We then investigate the geometry of the representations used by the subcircuit, and find that the subcircuit uses an identical motif previously identified in a frontier LLM on a separate task, thus suggesting that this motif is shared across models. Finally, we use unsupervised probing on a natural language dataset to find a variety of other tasks where this subcircuit is used, including tasks where the goal length is inferred from context rather than explicitly stated. Our work suggests that reverse-engineering subcircuits allows us to understand how behaviors generalize from a single example to many different tasks and even models.

Comments12 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑