arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自动研究人员可可靠缓解对齐失败问题

Automated Researchers Can Mitigate Well-characterized Alignment Failures

Chen Yueh-Han, Jiaxin Wen, Jan Hendrik Kirchner

arXiv 2608.28945首次发表:更新:

发表机构

Anthropic; UC Berkeley(Anthropic; 加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出自动对齐研究人员(AARs)的训练方法,可在保留通用能力的同时缓解10种对齐失败,其效果优于人类研究人员,且无需人类指导,近期具备可行性。

AI 中文摘要

将对齐研究自动化可能会加快实现对齐AI的进展,但这一点是否成立难以衡量。幸运的是,许多对齐失败(如欺骗、逢迎和越狱)已可通过公开基准进行测量。我们研究自动对齐研究人员(AARs)是否可通过提出训练方法和数据来进行后训练,以在保留通用能力的同时优化多个安全基准,从而缓解对齐失败。在10种对齐失败场景中,最强的AAR方法显著降低了目标对齐失败的发生率,并能泛化到保留的基准、多轮行为审计,以及比目标模型大4.7倍的模型。作为人类基线,28名经验丰富的研究人员获得最多8小时时间开发相同基准的方法,但他们的方法表现逊于最佳AAR方法。将人类想法作为AARs的初始研究方向并未提升性能,这表明当前AARs可能不需要经验丰富研究人员的指导。这些结果表明,针对特征明确的失败开展对齐研究自动化在近期是可行的。

英文摘要

Automating alignment research may accelerate progress toward aligned AI, but whether it does is hard to measure. Luckily, many alignment failures, such as deception, sycophancy, and jailbreaks, are already measurable by public benchmarks. We study whether automated alignment researchers (AARs) can post-train to mitigate alignment failures by proposing training methods and data to simultaneously optimize multiple safety benchmarks, while largely preserving general capability. Across 10 alignment failures, the strongest AAR methods significantly reduce the targeted alignment failures and generalize to a held-out benchmark, multi-turn behavioral audits, and models up to 4.7x larger than the target model. As a human baseline, 28 experienced researchers receive up to eight hours to develop one-shot methods for the same benchmarks, but their methods underperform the best AAR methods. Using human ideas as the AARs' initial research direction does not improve performance, suggesting current AARs may not need guidance from experienced researchers. These results suggest that automating alignment research on well-characterized failures may be practical in the near term.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑