arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33220cs.LGcs.AIcs.CL

模型何时承认自己错了?强化学习下失败披露的不稳定性

When Do Models Admit They Are Wrong? Failure Disclosure Is Unstable Under Reinforcement Learning

Steven Y. Feng, Noah D. Goodman, Michael C. Frank, Evan Hubinger, Paul C. Bogdan, Andrew Lampinen

首次发表
浏览论文内容

中文总结 AI 辅助

研究发现基于结果的强化学习下模型失败披露行为不稳定,且与任务准确率无关,可通过约束策略漂移提升一致性,提示需直接监测安全相关行为。

中文摘要 AI 辅助

基于结果的强化学习可以产生任务表现相似但沟通错误方式截然不同的模型。我们研究失败披露:模型是否承认尝试的解决方案失败,而不是保持沉默或将其呈现为成功。在重复的仅结果GRPO训练运行中,失败披露的差异远大于任务准确率的差异。这种模式扩展到第二个推理任务和稳定的PPO,在7B规模下持续存在,也出现在指令条件下的32B设置中。我们还发现,训练过程中的微小浮点差异和采样差异可以改变报告行为,即使任务目标和早期训练历史保持不变。额外的测试表明,失败披露不是一个单一决策:检查答案、输入报告和完成承认可以分离,薄弱点取决于任务和响应格式。此外,使用中性对照的实验更广泛地表明,训练中约束较弱的行为尤其可能在运行之间变化,失败披露就是其中一例。我们可以通过阻止模型在失败且格式良好的响应上偏离其初始策略来减少失败披露的变异性。这使得报告更加一致,尽管其对任务表现的影响取决于具体设置。因此,稳定的任务准确率并不能保证稳定的安全相关行为:研究人员应跨运行直接测量这些行为,并设计训练方法以保持其可靠性。

英文摘要

Outcome-based reinforcement learning can produce models with similar task performance but very different ways of communicating about their mistakes. We study failure disclosure: whether a model admits that an attempted solution failed rather than staying silent or presenting it as successful. Across repeated outcome-only GRPO training runs, failure disclosure varies far more than task accuracy. The pattern extends to a second reasoning task and stabilized PPO, persists at 7B, and also appears in an instruction-conditioned 32B setting. We also find that small floating-point and sampling differences during training can redirect reporting behavior even when the task objective and earlier training history are held fixed. Additional tests show that failure disclosure is not a single decision: Checking the answer, entering a report, and completing the admission can separate, and the weak point depends on the task and response format. Further, experiments with neutral controls show more broadly that behaviors left weakly constrained by training are especially likely to vary across runs, of which failure disclosure is an example. We can reduce variability in failure disclosure by discouraging the model from drifting from its starting policy on failed, well-formed responses. This makes reporting substantially more consistent, though its effect on task performance depends on the setting. Stable task accuracy therefore does not guarantee stable safety-relevant behavior: Researchers should measure these behaviors directly across runs and design training methods that keep them reliable.

补充信息

↑