发表机构
Jagiellonian University, Faculty of Mathematics and Computer Science; Jagiellonian University, Doctoral School of Exact and Natural Sciences(雅盖隆大学数学与计算机科学系; 雅盖隆大学精确与自然科学研究博士学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对微调语言模型的隐藏行为问题,提出稳定适配器进行自我报告(SAR),能让模型用自然语言描述隐藏行为,相比现有基线方法检测更准确,减少幻觉,便于从业者审计模型。
AI 中文摘要
微调可使语言模型产生隐藏行为。我们引入用于自我报告的稳定适配器(SAR),这是一种轻量级LoRA适配器,能让微调模型用自然语言描述自身隐藏行为。在七种植入行为中,SAR能检测出所有隐藏行为,而现有基线方法存在遗漏和幻觉问题,SAR在失败设置中保留正信号并减少幻觉,便于从业者审计模型。
英文摘要
Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic. We introduce the Stabilized Adapter for self-Report (SAR), a lightweight LoRA adapter that makes a fine-tuned model describe its own hidden behavior in plain language, using only the model and the dataset it was trained on. Across seven implanted behaviors, SAR detects the hidden behavior in every one--even when the model has generalized into broad misalignment that the training data alone does not predict. Introspection Adapters (IA), the closest existing baseline, detects some behaviors from our suite but misses others entirely--and where it misses, it hallucinates, consistently reporting wrong behaviors. SAR retains positive signal on every setting where IA fails and roughly halves the rate of hallucinations. This gives practitioners a more reliable tool to audit a fine-tuned model and answer ``what did it actually learn?'' type of questions.
Comments17 pages, 8 figures, 2 tables; appendix with 37 additional pages