arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可证明自身利用行为的智能体:用于安全利用对手的置信度调度受限响应

Agents That Certify Their Own Exploits: Confidence-Scheduled Restricted Responses for Safe Opponent Exploitation

Boning Li, Longbo Huang

arXiv 2607.28520首次发表:更新:

发表机构

IIIS, Tsinghua University(清华大学交叉信息研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出CS-RNR方法,通过置信序列跟踪对手行动频率,生成可自我审计的安全利用策略,在Leduc等博弈中实现远超基准的收益且符合预算要求。

AI 中文摘要

在双人零和不完全信息博弈中,采用纳什均衡策略的智能体能保证博弈价值,但会放弃有缺陷对手提供的额外价值。分散式偏离带来特殊挑战:二元释放规则可能收集到的证据过少而无法行动,而对不完整对手模型的完全最佳响应可能具有高度可利用性。我们提出了预算约束置信度调度受限响应(CS-RNR),这是首个对手利用方法,其安全保证是智能体对实际部署策略计算出的证明,因此智能体承诺的每一次利用都是自身审计过的。该方法用任意时间有效的置信序列跟踪汇总行动频率,仅当频率区间与均衡参考值分离时,才将其视为可利用的。经确认的偏离定义了保守对手模型,受限响应求解器将其转化为一系列 pin 级别网格上的候选反策略。部署前,每个完整候选策略都通过全树最佳响应进行评估。所得证明与用户指定的预算进行比较,并与策略原子性地提交。由于该检查是对所玩策略执行的,模型质量决定了实现的利用程度,而证明控制了相对于参考的预期损失。在 Leduc 扑克中,CS-RNR 获得了经金钱验证的二元门稳态收益的 6.2 倍,同时保持所有部署策略在预算内。使用相同估计器的轨迹混合达到了预算的 13.6 倍。在 Leduc、Liar's Dice 和 5 阶 Leduc 中,所有 36000 手经审计的牌都满足报告的证明容差。

英文摘要

An agent playing a Nash-equilibrium strategy in a two-player zero-sum imperfect-information game secures the game value but forfeits the additional value offered by a flawed opponent. Diffuse deviations pose a particular challenge: binary release rules may gather too little evidence to act, while a full best response to an incomplete opponent model can be highly exploitable. We introduce \emph{budget-constrained confidence-scheduled restricted responses} (CS-RNR), the first opponent-exploitation method whose safety guarantee is a certificate the agent computes on the strategy it actually deploys, so that every exploit it commits to is one it has audited itself. The method tracks pooled action frequencies with anytime-valid confidence sequences and treats a frequency as exploitable only once its interval separates from an equilibrium reference. The confirmed deviations define a conservative opponent model, which a restricted-response solve turns into candidate counter-strategies over a grid of pin levels. Before deployment, each complete candidate is evaluated by a full-tree best response. The resulting certificate is compared with a user-specified budget and committed atomically with the strategy. Because this check is performed on the played strategy, model quality determines the exploitation achieved while the certificate controls reference-relative expected loss. In Leduc hold'em, CS-RNR obtains $6.2\times$ the steady-state gain of a money-verified binary gate while keeping every deployed strategy within budget. A trajectory mixture using the same estimator reaches $13.6\times$ the budget. Across Leduc, Liar's Dice, and 5-rank Leduc, all $36{,}000$ audited hands satisfy the reported certificate tolerance.

Comments21 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑