发表机构
Institute of Science Tokyo; ZOZO Research; Stanford University; Waseda University; Nanyang Technological University(东京科学大学; ZOZO研究所; 斯坦福大学; 早稻田大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对微调引发的副作用对齐偏差问题,提出了Delta感知内省适配器(DAIA),构建了对应数据集,实验显示DAIA性能优于现有内省适配器且泛化性良好。
AI 中文摘要
微调可让源模型在目标领域获得所需能力与行为,同时保留其大部分通用能力,但这一适配过程可能会降低源模型原本具备的对齐属性。近期研究表明,大型语言模型可使用基于LoRA的模块(称为内省适配器,即IA)进行训练,以描述微调引发的行为变化。不过,现有研究主要关注模型在专门设计用于植入特定行为的数据集上进行微调,随后被要求解释该植入行为的场景,这与实际部署场景不同,实际部署场景的核心关注点往往是副作用对齐偏差:在与安全或对齐无明显关联的任务上进行微调,会导致对齐属性发生意外退化。为弥合这一差距,我们提出了一种名为“副作用内省”的新型问题设定,其中内省对象并非通过微调明确植入的行为,而是作为意外副作用出现的对齐偏移,并为该设定构建了数据集。此外,为提升对模型内部变化的敏感性,我们提出了Delta感知内省适配器(即DAIA),这是一种专门设计用于同时处理基础模型激活值和微调引发的激活值差异的新型机制。我们的实证评估表明,内省学习可泛化至未见过的微调模型和安全类别,且DAIA的性能始终优于现有内省适配器。
英文摘要
Fine-tuning enables a source model to acquire desired capabilities and behaviors in a target domain while retaining much of its general-purpose competence. However, this adaptation process can also degrade alignment properties that were present in the source model. Recent work has shown that large language models can be trained using LoRA-based modules known as introspection adapters (IAs) to describe behavioral changes induced by fine-tuning. However, existing studies primarily consider settings in which the model is fine-tuned on datasets explicitly designed to implant a specific behavior and is then asked to explain the implanted behavior. This differs from practical deployment scenarios, where the misalignment is not necessarily deliberately implanted, but arises only as a side effect. To bridge this gap, we formulate a novel problem setting, in which the target of introspection is not necessarily a behavior explicitly implanted through fine-tuning, but rather alignment shifts that may emerge as unintended side effects, and we construct a dataset for this setting. Furthermore, to enhance sensitivity to internal model changes, we propose the Delta-Aware Introspection Adapter (DAIA), a novel mechanism designed to explicitly process both base-model activations and activation differences induced by fine-tuning. Our empirical evaluation shows that introspection learning generalizes to unseen fine-tuned models and safety categories, and that DAIA generally outperforms existing IAs.