发表机构
Korea University; NAVER Cloud; KAIST(高丽大学; NAVER云; 韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对VL微调导致LLM文本能力损失的问题,提出用Sink Strength诊断并预测退化,发现QK-RMSNorm及现成权重合并无法恢复能力,建议用Sink Strength筛选主干。
AI 中文摘要
将预训练大语言模型(LLM)微调为视觉-语言模型(VLM)会削弱主干模型的文本能力,损伤集中在需要遵循精确输出规则的任务上,如指令遵循、基于严格解析最终答案的思维链推理以及带有严格评分者的类似评估。我们将这一差距归因于注意力汇聚点(attention-sink)的损坏:视觉-语言(VL)微调会扰动锚定大部分注意力概率的早期汇聚点位置,基础LLM对其汇聚点的保留程度决定了受影响能力在适配后保留多少。基于此观点,我们引入了Sink Strength,这是一个可在单GPU上几秒内基于基础LLM计算的标量,无需任何VL训练即可预测VL后的性能退化。它能一致地追踪6组VLM-LLM对及多个格式敏感任务的相对退化。作为该诊断的补充,我们发现预训练后注入的QK-RMSNorm无法重现原生QK-RMSNorm的保护作用,而几种现成的权重合并设置也无法恢复VL训练后丢失的能力。这些负面结果强调了在VL训练前用Sink Strength筛选主干模型的价值,并将干预方向缩小到头部选择性训练时保护。
英文摘要
Fine-tuning a pretrained LLM into a vision-language model (VLM) can erode the backbone's text capability, with the damage concentrated on tasks that require following exact output rules, such as instruction following, chain-of-thought reasoning graded on a strictly parsed final answer, and similar evaluations with strict graders. We trace this gap to attention-sink corruption: VL fine-tuning perturbs the early sink position that anchors a large fraction of attention probability, and how well the base LLM preserves its sink tracks how much of the affected capability survives adaptation. Building on this view, we introduce Sink Strength, a single scalar computed on the base LLM in a few seconds on a single GPU that predicts post-VL degradation without any VL training. It consistently tracks relative degradation across the six VLM-LLM pairs and multiple format-sensitive tasks. Complementing this diagnostic, we find that post-pretraining QK-RMSNorm injection fails to reproduce the protection of native QK-RMSNorm, while several off-the-shelf weight-merging settings fail to recover the lost capability after VL training. These negative results underscore the value of screening backbones with Sink Strength before VL training and narrow the intervention space toward head-selective training-time protection.
CommentsAccepted to EMNLP 2026. 29 pages, 6 figures, 15 tables. Code and data: https://github.com/minsik-choi126/sink-strength. * Equal contribution. † Corresponding author