arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

改进可扩展监督:协同训练的监控器

Improving scalable oversight with co-trained monitors

Joseph H. Rudoler, Kevin Tan, Benedict Tessler, Timothy Kong, Enric Boix Adserà

arXiv 2609.36049首次发表:更新:

发表机构

University of Pennsylvania(宾夕法尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出协同训练监控器以提升AI监督的可扩展性,证明监督式方法在有限Littlestone维度下可行,并引入测试时蒸馏的自监督方法,实验表明自适应监控器更能应对工人策略演变。

AI 中文摘要

工人-监控器设置是人工智能监督的一种有前景的方法,但针对固定监控器训练工人可能会激励监控器规避行为。我们研究是否可以通过与工人协同训练监控器来避免这种失败模式,并探索了监督式和自监督式两种方法。在监督式设置中,我们证明了一个特征:当且仅当可能的监控器函数类具有有限的Littlestone维度时,监控才能以消失的误差和查询率实现。这将工人监控与既有的对抗性在线学习文献联系起来。对于自监督,我们提出了一种基于测试时蒸馏的协同训练程序:监控器利用额外的测试时计算生成训练标签,然后在这些标签上训练其标准计算策略。对于多数投票标签,我们在具有动作覆盖的自适应工人分布下给出了有限样本锐化保证,表明监控器的判定收敛到其初始多数判定。我们在代码安全设置中对前者进行了压力测试,其中工人被对抗性训练以欺骗监控器。我们的结果表明,自适应监控器更能跟上不断演变的工人策略,而固定监控器更容易受到规避的影响。

英文摘要

Worker-monitor setups are a promising approach to AI oversight, but training workers against fixed monitors can incentivize monitor evasion. We study whether this failure mode can be avoided by co-training the monitor alongside the worker, and explore both supervised and self-supervised approaches. In the supervised setting, we prove a characterization: monitoring is possible with vanishing error and query rates exactly when the class of possible monitor functions has finite Littlestone dimension. This connects worker monitoring with an established literature on adversarial online learning. For self-supervision, we propose a co-training procedure based on test-time distillation: the monitor uses additional test-time compute to generate training labels, then trains its standard-compute policy on those labels. For majority-vote labels, we give a finite-sample sharpening guarantee under adaptive worker distributions with action coverage, that shows that the monitor's verdicts converge to its initial modal verdicts. We stress-test the former in code-security settings where the worker is trained adversarially to fool the monitor. Our results suggest that adaptive monitors are better at keeping pace with evolving worker strategies, while fixed monitors are more vulnerable to evasion.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑