arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过对探针进行训练实现对齐而不丧失可监控性

Alignment via Training Against Probes Without Losing Monitorability

Lena Libon, Alexander Panfilov, Ben Rank, Xin Chen, Jonas Geiping, Maksym Andriushchenko

arXiv 2609.38645首次发表:更新:

发表机构

ETH Zurich; ELLIS Institute Tübingen; MPI for Intelligent Systems; Tübingen AI Center(苏黎世联邦理工学院; ELLIS 图宾根研究所; 马克斯·普朗克智能系统研究所; 图宾根人工智能中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出探针引导微调,通过训练模型对抗持续更新的探针来增强无害性和诚实性,在保持效用的同时提升安全-效用权衡,并增强对越狱和消融攻击的鲁棒性,且不丧失可监控性。

AI 中文摘要

模型通常基于其观察到的输出,通过演示、偏好数据或奖励信号进行对齐。这些目标奖励看起来对齐的响应。更强大的模型可能学会满足这些目标而不内化预期行为,例如在训练期间假装合规。当目标定义在模型内部而非输出上时,这种表面合规可能更难实现。因此,我们研究了探针引导的微调,使用检测模型激活中不良属性的探针作为直接训练信号。我们在两个对齐目标(无害性和诚实性)上评估了线性和非线性探针,每层使用不同数量的探针。我们发现,针对训练期间不更新的探针进行训练是一个容易被利用的目标,而持续更新的探针在保持效用的同时显著降低了有害性并提高了诚实性。探针引导的微调在安全-效用权衡上优于DPO和推理时引导,同时对越狱和消融攻击的鲁棒性更强。此外,概念在微调后仍保持线性编码,这意味着我们的方法不会丧失监督。因此,针对探针进行训练提供了一种塑造模型所代表内容而非仅输出内容的方式,随着模型在使输出看起来对齐方面变得更好,这可能变得越来越重要。

英文摘要

Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals. These objectives reward responses that look aligned. More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance during training. Such superficial compliance could be harder when the objective is defined on model internals rather than outputs. Therefore, we study probe-guided fine-tuning, using probes that detect undesired properties in model activations as a direct training signal. We evaluate linear and non-linear probes with different numbers of probes per layer across two alignment objectives: harmlessness and honesty. We find that training against probes that do not update during training is an easily exploitable objective, while continuously updated probes substantially reduce harmfulness and improve honesty while preserving utility. Probe-guided fine-tuning achieves better safety-utility trade-offs than DPO and inference-time steering, while being substantially more robust against jailbreak and abliteration attacks. Moreover, the concepts stay linearly encoded after fine-tuning, meaning oversight is not lost by our method. Training against probes thus offers a way to shape what models represent rather than only what they output, which may become increasingly important as models get better at making their outputs look aligned.

Comments38 pages, 22 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑