发表机构
ServiceNow; University of Twente(ServiceNow(服务now公司); 特温特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对模型合并中解码器存在的表示偏差挑战,本文提出DARTS方法,通过熵加权L1损失与逐位置加性偏差校正,在三类任务上较标准方法获显著改进且新增参数极少。
AI 中文摘要
模型合并可将多个针对特定任务微调后的大型语言模型(LLMs)组合为单个多任务模型,无需额外训练,但合并后的模型存在表示偏差:即合并模型的隐藏状态与各源模型隐藏状态间存在系统性漂移。Yang等人(2024a)针对基于编码器的视觉模型,使用经L1损失训练的轻量校正模块研究并缓解了该偏差。然而,由于解码器的自回归特性,此前未针对解码器模型研究此类偏差。本文分析了解码器模型的表示偏差问题,发现了编码器不存在的两项挑战:(1)因果注意力掩码会导致偏差在各标记位置累积,需采用依赖位置的校正;(2)并非所有标记位置同等重要,高熵(决策关键)位置远比低熵位置重要。为应对这些挑战,本文提出了感知解码器的表示调优方法(DARTS)。DARTS采用新颖的熵加权L1损失,对高熵位置的校正赋予更高权重,此类位置的误差对生成质量影响最大;同时采用逐位置加性偏差,可捕捉依赖位置的误差且不会过度参数化。本文在三个领域开展了广泛评估:代码生成(HumanEval)、数学推理(GSM8K)和指令遵循(AlpacaEval),使用Llama-2-7B模型,结果显示,DARTS相较标准“手术”方法实现了显著改进,且新增参数可忽略(仅占总参数的0.1%)。
英文摘要
Model merging combines multiple task-specific fine-tuned LLMs into a single multi-task model without additional training. However, merged models are known to suffer from representation bias: systematic drift between the merged model's hidden states and those of each individual source model. Prior work (Yang et al., 2024a) study and mitigate this bias for encoder-based vision models using a lightweight correction module trained with L1 loss. However, such bias is not studied for decoder models due to their autoregressive nature. We analyze the problem of representation bias in decoder models, and show two challenges absent in encoders: (1) the causal attention mask causes bias to accumulate across token positions, requiring position-dependent correction; and (2) not all token positions are equally important, i.e., high-entropy (decision-critical) positions matter far more than low-entropy ones. To address these challenges, we propose Decoder-Aware Representation Tuning via Surgery (DARTS). DARTS employs a novel entropy-weighted L1 loss to upweight correction at high-entropy positions where errors most affect generation quality, and a per-position additive bias that captures position-dependent error without overparameterization. We perform extensive evaluation on three domains: code generation (HumanEval), mathematical reasoning (GSM8K), and instruction following (AlpacaEval) on Llama-2-7B models, and show DARTS achieves significant improvement over the standard surgery approach while adding negligible parameters ($0.1\%$ of total parameters).
CommentsAccepted to EMNLP 2026 Main Conference