如何训练你的模式生物
How to train your model organism
浏览论文内容
中文总结 AI 辅助
本文提出用三个目标(行为安装、能力保持、输出自然度)验证模式生物,并基于模型合并的多目标训练方法提升其逼真度,以确保可解释性结论的可靠性。
中文摘要 AI 辅助
对齐相关行为(如后门、谄媚、虚假关联)的模式生物已成为评估白盒可解释性技术的关键工具。我们认为,当前将模式生物训练到仅安装目标行为这一单一目标的做法是不够的,并建议根据三个目标及其相关指标来验证模式生物:目标行为安装、通用能力保持(即参数知识、聊天质量)以及输出自然度(即思维链和激活)。我们使用这一验证框架重新审视了两个公开发布的生物套件,并表明:(1)聊天质量和思维链自然度在各种训练方案中显著下降;(2)验证指标能预测可解释性方法恢复已安装行为的程度,例如,logit透镜读出与生物的通用能力共变。我们引入了一种基于模型合并的多目标训练方法,以训练更逼真的模式生物。最后,在一个针对临床推理中人口统计偏见的新模式生物套件上,我们比较了训练方案,发现DPO训练比监督微调更接近基础模型,且所提出的模型优化方法能更好地保持能力和自然度。使用调查代理审计该套件时,我们再次观察到验证指标追踪偏见恢复情况。总之,训练方法塑造了生物所支持的可解释性结论,我们认为应考虑多个目标,以使用(逼真的)模式生物得出关于可解释性方法的可推广结论。
英文摘要
Model organisms of alignment-relevant behaviors (e.g., backdoors, sycophancy, spurious correlations) have emerged as a key tool for evaluating whitebox interpretability techniques. We argue that the prevailing practice of training model organisms to a single objective of installing the target behavior is insufficient and propose validating model organisms with respect to three objectives with associated metrics: target-behavior installation, general-capability preservation (i.e., parametric knowledge, chat quality), and output naturalness (i.e., CoT and activations). We re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities. We introduce a multi-objective training approach based on model merging to train more realistic model organisms. Finally, on a new suite of model organisms targeting demographic biases in clinical reasoning, we compare training recipes and find that DPO training stays closer to the base model than supervised finetuning, and the proposed model optimization approach better preserves capabilities and naturalness. Auditing this suite with an investigator agent, we again observe validation metrics tracking bias recovery. In sum, training methods shape the interpretability conclusions an organism supports, and we argue that one should consider multiple objectives to draw generalizable conclusions about interpretability methods using (realistic) model organisms.
发表机构
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。