Where's the Plan? Locating Latent Planning in Language Models with Lightweight Mechanistic Interventions
计划在哪里?通过轻量级机制干预定位语言模型中的潜在规划
机构 * University of California, Berkeley(加州大学伯克利分校)
专题命中 知识编辑与模型理解 :language model(title,abstract);分类 cs.AI、cs.LG
AI总结 通过押韵对句补全任务,使用线性探针和激活修补方法,研究语言模型在生成过程中是否形成并因果依赖未来约束的潜在规划,发现仅Gemma-3-27B模型存在因果依赖,并定位到五个注意力头。
Comments 13 pages, 20 figures, 3 tables. Accepted to Workshop on Mechanistic Interpretability @ ICML 2026