发表机构
University of Liverpool; Amazon(利物浦大学; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究探讨微调如何改变LLM内部表征,发现EAP识别组件集中于特定层且与表征变化层不相关,且任务间组件重叠不带来性能迁移,甚至可能相互损害。
AI 中文摘要
微调已成为将大型语言模型(LLM)适配到各种下游任务的广泛采用的方法。然而,微调如何重塑其内部机制仍知之甚少。为解决这一问题,我们研究了微调如何改变LLM中的内部表征,包括注意力模式和逐层激活,并考察这些变化是否与EAP(基于激活的因果路径)识别的任务相关组件(如注意力头和logit级激活)相关,这些组件驱动任务性能。我们发现,EAP识别的组件集中在特定层内,表明模型内化任务特定行为时存在一定程度的功能局部化。值得注意的是,这些组件在层间的分布与微调期间经历最显著表征变化的层在很大程度上不相关。此外,我们观察到,如果任务性质不同(例如分类任务与生成任务),EAP识别组件在不同任务间的重叠并不会转化为跨任务性能迁移。更具体地说,当两个任务在其EAP识别组件上表现出高度重叠时,在一个任务上的微调可能导致在另一个任务上的性能下降。
英文摘要
Fine-tuning has emerged as a widely adopted approach for adapting LLMs to a variety of downstream tasks. However, how it reshapes their internal mechanisms remains poorly understood. To address this, we investigate how fine-tuning alters internal representations in LLMs, including attention patterns and layer-wise activations, and examine whether these changes are linked to task-relevant components identified by EAP (e.g., attention heads and logit-level activations) that drive task performance. We find that EAP-identified components are concentrated within specific layers, indicating a degree of functional localisation in how models internalise task-specific behavior. Notably, the distribution of these components across layers is largely uncorrelated with the layers undergoing the most substantial representational changes during fine-tuning. Furthermore, we observe that overlap in EAP-identified components across tasks does not translate into cross-task performance transfer if the tasks are different in nature (e.g. classification vs. generative tasks). More specifically, fine-tuning on one task can lead to a degradation of performance on another when the two tasks exhibit a high degree of overlap in their EAP-identified components.
Comments25 pages, 14 figures, 7 tables. Accepted at AACL-IJCNLP 2026