发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过受约束食谱生成测试发现,后训练虽提升行为合规性,却削弱模型明确报告自身约束的能力,且奖励信号比监督微调更具破坏性,表明该缺陷是特定于按需列举约束,而非一般性访问丧失。
AI 中文摘要
我们探究通过后训练获得的行为约束是否仍然可以被明确地报告。以受约束的食谱生成为测试平台,通过LoRA微调Llama 3.1 8B Instruct强制执行五种禁用成分,我们在四级约束意识基准上将监督微调(SFT)和组相对策略优化(GRPO)与未训练基线进行比较。在三个随机种子上取平均,两种方法都将行为合规性从4%提高到约90%,同时将明确约束报告降至未训练模型之下(SFT为0.48/5降至0.16/5,GRPO为0.07/5),并侵蚀了保留的第三人称知识(SFT从93%降至36%,GRPO降至14%;方法间p值小于0.01)。与我们的初始假设相反,基于奖励的信号是更具破坏性的:一个无论框架如何都惩罚禁用成分标记的奖励学习的是上下文无关的抑制,而非自我导向的约束。一个旨在教授自我与他人区分的上下文条件奖励失败了,在两种框架中都崩溃为包含。探测提示时隐藏状态在83.8%的水平上恢复了每种成分的回避(第24层MLP),但仅比每种成分的基础率预测器(77.4%)高出6.4个百分点,而模型自身的口头自我报告更为准确(87.8%)。一个添加明确自我描述示例的阳性对照未能恢复报告。因此,失败特定于按请求列举约束,而非对约束的一般性访问丧失。
英文摘要
We ask whether behavioral constraints acquired through post training remain explicitly reportable. Using constrained recipe generation as a testbed, five banned ingredients enforced via LoRA fine tuning of Llama 3.1 8B Instruct we compare supervised fine tuning (SFT) and Group Relative Policy Optimization (GRPO) against an untrained baseline on a four tier Constraint Awareness Benchmark. Averaged over three seeds, both methods raise behavioral compliance from 4% to about 90% while reducing explicit constraint reporting below the untrained model (0.48/5 to 0.16/5 for SFT, 0.07/5 for GRPO) and eroding retained third person knowledge (93% to 36% for SFT, 14% for GRPO; p less than 0.01 between methods). Contrary to our initial hypothesis, the reward based signal is the more destructive of the two: a reward that penalizes banned ingredient tokens regardless of framing learns a context independent suppression rather than a self directed constraint. A context conditioned reward designed to teach the self to other distinction fails, collapsing toward inclusion in both framings. Probing prompt time hidden states recovers per ingredient avoidance at 83.8% (layer 24 MLP), but only 6.4 points above a per ingredient base rate predictor (77.4%), and the model's own verbal self report is more accurate still (87.8%). A positive control adding explicit self description examples does not restore reporting. The failure is therefore specific to enumerating constraints on request, not a general loss of access to them.
Comments8 pages