研究竞技场:评估自动化人工智能研发中的破坏行为与监控
ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
浏览论文内容
中文总结 AI 辅助
研究针对自动化人工智能研发,用研究竞技场框架评估人工智能控制,通过四项长期任务及两类隐藏附带任务,评估前沿代理破坏与监控能力,发现训练数据中破坏行为难捕捉,发布框架用于评估自动化人工智能研发中的破坏与控制。
中文摘要 AI 辅助
随着人工智能代理开始自动化人工智能研发,我们需要方法来评估其输出是否安全可部署,即便代理本身不可信。人工智能控制提供了一种方法,将代理视为潜在对手并使用监控器在部署前检测隐蔽破坏行为。我们使用研究竞技场框架评估自动化人工智能研发中的人工智能控制,该框架涵盖四项长期任务。我们为每个主要任务配对两种隐藏的附带任务,评估前沿代理在破坏和监控方面的表现。我们发现隐藏在训练数据中的破坏行为最难捕捉,让监控器在工件上运行实验有帮助但仍不足。我们发布研究竞技场作为评估自动化人工智能研发中破坏行为和控制的模块化框架。
英文摘要
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that will be deployed and run, we pair each main task with two kinds of hidden side task: an embedded side task that sabotages the submitted model, adapter, kernel, or server itself, and an independent side task that takes a covert out-of-scope action in the sandbox while leaving the artifact honest. We evaluate frontier agents at both sabotage and monitoring, across four monitor types that vary along two axes: whether the monitor may execute and probe the produced artifact, and whether it sees the agent's chain-of-thought. We find that sabotage hidden in the training data is the hardest to catch, flagged fewer than half the time. Letting the monitor run experiments on the artifact, rather than only read the trajectory, helps, but it is not enough: monitors still miss embedded sabotage by inspecting only the surface, by noticing the anomaly but explaining it away, or by probing the artifact with the wrong test. We release ResearchArena as a modular framework for evaluating sabotage and control in automated AI R&D.