arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09002cs.AIcs.MA

在多智能体博弈中学习报告不安全任务

Learning to Report Unsafe Tasks in a Multi-Agent Game

发表机构密歇根大学法学院 · 新加坡国立大学
查看机构详情
  • University of Michigan Law School(密歇根大学法学院)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Avyay M. Casheekar, Hariganesh Tangirala

首次发表
浏览论文内容

中文总结 AI 辅助

本研究在多智能体博弈中探讨如何通过审计促使智能体报告不安全任务,提出审计条件并训练PPO策略,证明分离目击者能显著降低不安全完成率,而完全共享策略在独立采样下难以达标。

中文摘要 AI 辅助

当智能体因完成任务而共享奖励时,报告不安全工作可能通过停止任务来减少报告者的奖励。审计可以使报告成为最优选择,但并不能确保进一步的训练教会沉默的团队进行报告。我们在一个任何目击者都可以通过报告来停止任务的博弈中研究这一学习问题。每个任务有$k$个共享策略且独立采样的目击者,在普遍沉默时,期望奖励对共享沉默概率的导数会将每个任务的收益计数$k$次。与普遍报告的比较则只计数一次。对于任意策略组,我们给出了一个审计条件,该条件足以使精确的策略梯度更新达到普遍报告,并且在边界情况之外,在普遍沉默附近是必要的。在一个平衡族中,满足该条件并具有规定正边际的最便宜审计,在完全共享时比每个角色一个策略时恰好贵$k$倍。我们在24个目击者图上从学习到的沉默中训练PPO策略。将共同目击者分开,与相同规模下经过相同审计的洗牌组相比,不安全完成率降低了33.59个百分点(95%图自举区间:21.03-45.13)。在48个目击者组运行中,只有9个在保持至少90%合法完成率的同时实现了低于1%的不安全完成率。在相同的审计预算下,一个完全共享的网络在独立动作采样的48次运行中没有一次达到这两个阈值,而在共同采样时48次全部达到。

英文摘要

When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task. Audits can make reporting optimal without ensuring that further training teaches a silent team to report. We study this learning problem in a game where any witness can stop a task by reporting. With $k$ witnesses per task sharing a policy and drawing independently, the expected-reward derivative with respect to their shared silence probability counts each task's benefit $k$ times at universal silence. The comparison with universal reporting counts it once. For arbitrary policy groups, we give an audit condition sufficient for exact policy-gradient updates to reach universal reporting and, apart from boundary cases, necessary near universal silence. In a balanced family, the cheapest audits meeting the condition with prescribed positive margins cost exactly $k$ times as much for full sharing as for one policy per role. We train PPO policies on 24 witness graphs from learned silence. Separating co-witnesses reduces unsafe completion by 33.59 percentage points compared with shuffled groups of the same sizes under the same audits (95% graph-bootstrap interval: 21.03-45.13). Only 9 of 48 witness-group runs achieve below 1% unsafe completion while retaining at least 90% legitimate completion. At the same audit budget, a fully shared network meets both thresholds in none of 48 runs with independent action draws and all 48 with a common draw.

↑