arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36221cs.LG

语言模型在串行需求下如何差异性地重新分配注意力头活动

How Language Models Differ in Redistributing Attention-Head Activity Under Serial Demand

Johnny Jingze Li, Abdulla Kuleib, Kalyan Basu, Gabriel A. Silva

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过固定提示长度并改变串行步骤,测量17个开放权重模型中注意力头的集中与分散模式,发现模型间存在可复现差异,且集中度高的模型更依赖关键头,为模型比较提供了新视角。

中文摘要 AI 辅助

模型在每一层注意力头上的活动分布方式,提供了其如何通过深度路由信息的粗略视图;这种分布如何随任务变化,是机制性解释必须阐明的一部分。在保持提示长度固定的情况下,我们改变任务所需的串行步骤数量,并测量17个开放权重模型的每一层中,随着需求增加,活动是集中在少数几个头上还是分散到多个头上。两种情况都会发生:在大多数模型中,深度中段之前的层会集中活动,而较深的层则分散活动。模型之间在发生位置和强度上存在差异。例如,Qwen2.5基础模型(从0.5B到7B)在每项任务中的分散程度低于平均模型,并在其后半部分的某些区域集中活动,而Llama模型(从1B到8B)和OLMo-2则在这些区域分散;这种对比在很大程度上也适用于层数和头数相同的Llama-3.1-70B与Qwen2.5-72B。这些差异是可复现的,且经过后训练的模型在很大程度上保留了其基础模型的模式。一项消融研究表明,在任务内,活动更集中于顶部头的模型,其答案也更多地依赖这些头。因此,跨层的集中与分散提供了一种新的模型比较方式,即通过它们如何通过深度路由信息来比较。代码可在以下网址获取:https URL。

英文摘要

The way a model distributes activity over each layer's attention heads offers a coarse view of how it routes information through depth; how this changes with the task is part of what a mechanistic account must explain. Holding prompt length fixed, we vary how many serial steps a task demands and measure, in every layer of 17 open-weight models, whether activity concentrates on a few heads or spreads across many as demand rises. Both occur: in most models, layers just before mid-depth concentrate activity and later layers spread it. Models differ in where and how strongly this happens. The Qwen2.5 base models from 0.5B to 7B, for example, spread less than the average model in every task and concentrate activity in parts of their second half, where Llama models from 1B to 8B and OLMo-2 spread; the contrast largely holds between Llama-3.1-70B and Qwen2.5-72B, which have the same number of layers and heads. These differences are reproducible, and post-trained models keep much of their base model's pattern. An ablation study suggests that, within a task, models whose activity is more concentrated on their top heads also depend more on those heads for the answer. Concentration and spreading across layers thus offer a new way to compare models, by how they route information through depth. Code is available at https://github.com/johnnyjli/serial-demand-heads.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • University of California, San Diego(加州大学圣迭戈分校)
  • Qualtrics LLC(Qualtrics 有限责任公司)

机构由 AI 辅助整理,请以论文原文为准。

↑