LLM组合任务中的因果与可解释结构
Causal and Interpretable Structures in LLM Compositional Tasks
浏览论文内容
中文总结 AI 辅助
本研究通过分析循环概念任务中的激活,揭示了LLM在Transformer层中逐步组织关系信息的几何与因果机制,并发现限制到因果相关联合表示可提升预测准确性。
中文摘要 AI 辅助
大型语言模型能够解决其答案不仅依赖于单个输入标记,还依赖于标记之间关系的任务。这种关系信息是如何在Transformer层中被表示和处理的?我们研究了来自提示集合的激活,这些提示要求推断对应于循环概念(月份、小时、星期几和音符)的三个标记之间的关系,以正确预测下一个标记。在模型家族(Llama、Qwen、Gemma和Mistral)和循环概念中,我们发现了标记间联合依赖在几何组织和因果使用上的一致逐层进展:中间层使用基于推断的两个标记之间关系的联合表示,而后续层使用与所有三个标记相关的联合表示来正确完成任务。我们还发现了标记之间的其他关系,这些关系在几何上结构化但在下一个标记预测中因果上不活跃。关键的是,当将这些几何和因果调查放在一起时,它们揭示了表示级别的机制,该机制逐步组织和组合关系信息以形成答案。更令人惊讶的是,将模型限制在这种因果相关的联合表示上可以提高下一个标记预测的准确性。
英文摘要
Large language models are able to solve tasks whose answers depend on not only individual input tokens, but also on relations among them. How is such relational information represented and processed across transformer layers? We study activations from ensembles of prompts that require inferring relationships between three tokens corresponding to a cyclic concept (months, hours, weekdays, and musical notes) to correctly predict the next token. Across model families (Llama, Qwen, Gemma, and Mistral) and cyclic concepts, we find a consistent layerwise progression in how the joint dependence among the tokens is geometrically organized and causally used: intermediate layers use a joint representation based on the inferred relationship between two tokens, while later layers use a joint representation associated with all three tokens to correctly complete the task. We also find other relationships between tokens that are geometrically structured but remain causally inert in the next-token prediction. Crucially, when taken together, these geometric and causal investigations reveal the representation-level mechanism that progressively organizes and composes the relational information to form the answer. More surprisingly, restricting the models to such causally relevant joint representations improves next-token prediction accuracy.
发表机构
- Cornell University(康奈尔大学)
- Goodfire AI
机构由 AI 辅助整理,请以论文原文为准。