大语言模型任务适应如何重塑对齐:行为和表征漂移的多维研究
How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift
浏览论文内容
中文总结 AI 辅助
研究大语言模型任务适应对对齐的影响,通过系统评估监督微调、KL正则化SFT和可验证奖励强化学习等方法,发现训练后调整非均匀重塑对齐,不同方法影响各异,强调任务适应是对齐干预,应进行多维对齐评估。
中文摘要 AI 辅助
训练后调整是使大语言模型适应下游任务的关键机制。先前研究表明任务适应会改变模型原有的对齐方式,但其在对齐领域的更广泛影响尚不清楚。我们通过对代表性任务适应方法进行系统评估来填补这一空白,涵盖监督微调(SFT)、KL正则化SFT和可验证奖励强化学习(RLVR),涉及安全、事实性等六个关键领域的15个对齐方面。结果表明训练后调整并非均匀重塑对齐。RLVR提高任务性能且导致较小但非零的特定指标偏移,SFT会导致跨领域更大对齐漂移,KL正则化减轻了这种影响。表征层面分析支持此模式。这些结果表明任务适应不仅是提升能力的步骤,本身也是一种对齐干预,促使多维对齐评估成为训练后管道的标准组成部分。
英文摘要
Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.
发表机构
- University of Cambridge(剑桥大学)
机构由 AI 辅助整理,请以论文原文为准。