arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02632cs.LG

大型语言模型(LLMs)可对归因图进行标注

LLMs Can Annotate Attribution Graphs

Ameen Patel, Max Zhang, Nathan Hu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出一种LLM自动化流程,可将特征分组为超级节点标注归因图,其效果与人类标注相当,在两跳首都任务中准确率达97%,还可用于开放式探索,推动自动化电路追踪研究。

中文摘要 AI 辅助

电路追踪是揭示语言模型内部计算的一项令人兴奋的技术,但它需要耗时的手动步骤,即把单个特征或多层感知机(MLP)神经元分组为超级节点。本文提出了一种自动化该步骤的简单流程:直接向语言模型提供特征描述,由其将特征分组为超级节点。通过自动化可解释性指标,我们确认该流程生成的超级节点与人类标注者生成的超级节点具有同等可解释性。在两跳首都任务中,该流程在100个提示中,有97个成功恢复了对应中间跳的超级节点。最后,我们展示了该流程用于开放式探索的简单概念验证:自动标注来自维基百科提示补全的1000个归因图,再利用大型语言模型(LLM)评判器标记值得人类审查的有趣图。我们希望这项工作表明,即使是简单的自动化也能生成有意义的归因图标注,推动自动化电路追踪领域的进一步研究。

英文摘要

Circuit tracing is an exciting technique for revealing the internal computation of language models, but it requires a time-intensive manual step of grouping individual features or MLP neurons into supernodes. We present a simple pipeline for automating this step: directly presenting feature descriptions to a language model that groups them into supernodes. Using automated interpretability metrics, we confirm that supernodes generated by our pipeline are as interpretable as those generated by human annotators. On a two-hop Capitals task, our pipeline recovers a supernode corresponding to the intermediate hop in 97 of 100 prompts. Finally, we present a simple proof of concept using our pipeline for open-ended exploration, where we automatically annotate 1000 attribution graphs from Wikipedia prompt completions and then use an LLM judge to flag interesting graphs worth human review. We hope this work demonstrates that even simple automation can produce meaningful attribution graph annotations, motivating further work on automated circuit tracing.

补充信息

↑