动态位置注意力调制用于大型语言模型的参数高效微调
Dynamic Positional Attention Modulation for Parameter-Efficient Fine-Tuning of Large Language Models
浏览论文内容
中文总结 AI 辅助
针对现有PEFT方法均匀静态适应的问题,提出DyPAM,通过维度、头、层调制位置注意力,在数学和常识推理基准上优于强基线。
中文摘要 AI 辅助
参数高效微调(PEFT)已成为将大型语言模型适应下游任务的标准方法。然而,大多数现有的PEFT方法依赖于均匀且静态的适应,未考虑注意力在维度、头、层和输入令牌之间的结构化异质性。在实践中,注意力表示表现出非均匀行为,而诸如旋转位置嵌入(RoPE)之类的位置编码机制会引发依赖于维度的位置结构,使得均匀适应效果欠佳。在这项工作中,我们提出了DyPAM(动态位置注意力调制),一种PEFT方法,通过直接作用于查询和键表示来调整位置信息对注意力的贡献。DyPAM结合了输入条件化的维度级调制与头级和层级结构调制,在不修改预训练骨干的情况下,对与RoPE诱导结构对齐的位置注意力进行细粒度适应。在多个骨干模型上的数学和常识推理基准的广泛实验表明,DyPAM始终优于现有的强PEFT基线。
英文摘要
Parameter-efficient fine-tuning (PEFT) has become a standard approach for adapting large language models to downstream tasks. However, most existing PEFT methods rely on uniform and static adaptations, without accounting for the structured heterogeneity of attention across dimensions, heads, layers, and input tokens. In practice, attention representations exhibit non-uniform behavior, and positional encoding mechanisms such as rotary positional embeddings (RoPE) induce dimension-dependent positional structure, making uniform adaptation suboptimal. In this work, we propose DyPAM (Dynamic Positional Attention Modulation), a PEFT method that adapts how positional information contributes to attention by operating directly on the query and key representations. DyPAM combines input-conditioned, dimension-wise modulation with head-wise and layer-wise structural modulation, performing fine-grained adaptation of positional attention aligned with the RoPE-induced structure without modifying the pretrained backbone. Extensive experiments on mathematical and commonsense reasoning benchmarks across multiple backbone models demonstrate that DyPAM consistently outperforms existing strong PEFT baselines.
发表机构
- Beihang University(北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。