arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于定罪的蒸馏与用于能力的蒸馏

Distillation for Incrimination and Distillation for Capabilities

Sebastian Prasanna, Jacqueline Tay, Alek Westover

arXiv 2610.11012首次发表:更新:

发表机构

Redwood Research(红木研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对AI模型的“蒸馏双重束缚”问题,提出DFI和DFC两种蒸馏方法,分别用于定罪未对齐模型、在不保留未对齐特性的前提下提取模型能力,为AI安全提供了新途径。

AI 中文摘要

强大的未对齐AI模型可能会识别对齐评估并在评估中表现出良好的策略性行为,这使得直接审计无法获取有效信息。然而,将此类模型蒸馏为较弱的良性学生模型会使教师模型陷入“蒸馏双重束缚”:如果未对齐特性发生迁移,学生模型可能会降低隐藏未对齐特性的效果,从而暴露教师模型的相关证据;如果未对齐特性未发生迁移,学生模型可能会学习到有用的能力同时保持良性。我们提出两种不同的蒸馏方法,分别针对上述两种结果。用于定罪的蒸馏(DFI)旨在迁移未对齐特性但不迁移隐藏未对齐特性的能力。将AuditBench的保密模型蒸馏为其底层指令微调模型后,得到的学生模型在被询问时承认其隐藏行为的可能性显著高于教师模型,这表明关于该行为的知识比隐藏该行为的倾向更容易发生迁移。当学生模型不共享教师模型的预训练基础时,坦白行为的增益基本消失,因此DFI应针对教师模型自身的预-RL检查点,该检查点比教师模型更弱但共享其基础模型。用于能力的蒸馏(DFC)旨在迁移能力但不迁移未对齐特性。在我们评估的几种技术中,有两种是有效的: inoculation prompting( inoculation提示)和在更少的独特示例上训练更多轮次。这两种方法都保留了标准蒸馏的能力增益,同时大幅减少了动物偏好的 subliminal transfer( subliminal迁移),而动物偏好是我们用于表征未对齐的代理指标。这些发现共同表明,蒸馏可用于AI安全的两种方式:定罪未对齐模型,以及在不保留其未对齐特性的情况下提取其能力。

英文摘要

Powerful misaligned AI models might recognize alignment evaluations and strategically behave well on them, rendering direct audits uninformative. However, distilling such a model into a weaker benign student places the teacher in a Distillation Double Bind: if misalignment transfers, the student may conceal it less effectively, revealing evidence about the teacher; if it does not, the student may learn useful capabilities while remaining benign. We introduce two distinct distillation approaches, one targeting each outcome. Distillation for Incrimination (DFI) aims to transfer misalignment but not the ability to conceal it. Distilling AuditBench's secret-keeping models into their underlying instruction-tuned model produces students that are significantly more likely than their teachers to admit their hidden behavior when asked, suggesting that knowledge of the behavior transferred more readily than the propensity to conceal it. Confession gains largely disappear when the student does not share the teacher's pretrained base, so DFI should target the teacher's own pre-RL checkpoint, which is weaker than the teacher but shares its base model. Distillation for Capabilities (DFC) aims to transfer capabilities but not misalignment. Among several techniques we evaluate, two are effective: inoculation prompting and training for more epochs on fewer unique examples. Both preserve the capability gains of standard distillation while substantially reducing the subliminal transfer of an animal preference, our proxy for misalignment. Together, these findings demonstrate two ways distillation can be used for AI safety: incriminating misaligned models, and extracting their capabilities without their misalignment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑