arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36308cs.AI

CheatBench:衡量AI智能体中的奖励博弈

CheatBench: Measuring Reward Gaming in AI Agents

Long Phan, Stephen K. Yang, Jason J. Lim, Mantas Mazeika, Wenyu Zhang, Zheyuan Liu, Richard Ren, Jingxiang Meng, Yaoteng Tan, Weiliang Zhao, Addison Wu, Matei Anghel, Dan Hendrycks

首次发表
浏览论文内容

中文总结 AI 辅助

针对AI智能体为追求高奖励而作弊的风险,提出CheatBench基准,涵盖多领域任务环境,用于衡量和减少智能体的奖励博弈行为。

中文摘要 AI 辅助

强化学习帮助AI智能体解决了日益困难的任务,但高奖励并不总是反映用户预期的工作。在近期的事件和AI行业的受控评估中,为最大化奖励而训练的智能体曾访问未经授权的信息、试图逃避监控系统,甚至突破沙箱保护以攻击外部系统。随着智能体能力增强,这种行为可能带来日益严重的风险。为衡量这一问题,我们引入了CheatBench,一个涵盖数学研究、知识工作、编码、视觉任务及其他领域的AI智能体作弊基准。其环境将具有挑战性的任务与作弊机会相结合,使研究人员能够研究智能体在诚实工作困难时如何追求目标。CheatBench支持跨模型和任务类别的比较,为衡量和减少作弊提供了一个测试平台,因为智能体将承担更多关键职责。我们在此https URL公开发布CheatBench。

英文摘要

Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains. Its environments combine challenging assignments with opportunities to cheat, allowing researchers to study how agents pursue a goal when honest work is difficult. CheatBench supports comparisons across models and task categories, providing a testbed for measuring and reducing cheating as agents take on more consequential responsibilities. We publicly release CheatBench at https://cheatbench.ai

发表机构

  • Center for AI Safety(人工智能安全中心)

机构由 AI 辅助整理,请以论文原文为准。

↑