arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过强化学习训练对齐审计员

Training Alignment Auditors via Reinforcement Learning

Paul Rosu, Rowan Wang

arXiv 2608.25460首次发表:更新:

发表机构

Anthropic(Anthropic)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究用强化学习训练LLM审计员,通过成对奖励等优化提升审计质量与真实性,误报率低于1%,且审计能力可跨框架泛化。

AI 中文摘要

前沿模型的对齐审计越来越依赖大语言模型(LLM)审计员来大规模发现不良行为,但当前自动化审计员在连贯调查和审计真实性方面存在困难。本研究用强化学习改进LLM审计员:在最优训练环境中,策略调查可能通过系统提示植入隐藏行为的目标模型;LLM评判员知晓目标是否存在隐藏行为,将策略的调查与参考调查进行整体比较以确定奖励。系统性消融实验显示,与逐点奖励相比,成对奖励能实现更稳健的训练,加入无植入行为的目标有助于维持低误报率。训练提升了对含植入行为目标的调查质量、在未修改生产模型中发现不良行为的比例及审计真实性,误报率保持在1%以下;此外,审计能力可跨框架泛化,在AuditBench的对抗微调目标上的性能显著提升(Sheshadri等人,2026)。

英文摘要

Alignment auditing of frontier models increasingly relies on LLM auditors to surface undesirable behaviors at scale, but current automated auditors can struggle with coherent investigation and audit realism. In this work, we improve LLM auditors with reinforcement learning. In our best training environment, the policy investigates target models that potentially possess hidden behaviors planted via their system prompt. An LLM judge, which knows whether the target has a hidden behavior, holistically compares the policy's investigation to a reference investigation to determine the reward. With systematic ablations, we find that pairwise rewards yield more robust training compared to pointwise rewards, and that adding targets without planted behaviors helps maintain a low false positive rate. Training improves investigation quality against targets with planted behaviors, the rate of concerning behaviors surfaced in unmodified production models, and audit realism, while false-positive rates stay below 1%. Furthermore, auditing capabilities generalize across scaffolds: performance on AuditBench's adversarially fine-tuned targets substantially improves [Sheshadri et al., 2026].

Comments82 pages, 15 figures. Code, prompts, and evaluation data are available at https://github.com/paulrosu11/training-auditing-agents-public

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑