基于扰动遗憾最小化的带水印博弈求解
Watermarked Game Solving via Perturbed Regret Minimization
浏览论文内容
中文总结 AI 辅助
本文针对现有博弈智能体水印技术仅适用于完美信息博弈的局限,提出集成于学习过程、可用于不完美信息博弈的扰动遗憾最小化水印方法,其水印代价小且易检测。
中文摘要 AI 辅助
博弈论可建模自利主体间的诸多现实交互,AI的快速发展引发了对不良主体滥用超人类或人类水平博弈智能体的担忧,包括意外或故意滥用。AI水印主要应用于大语言模型(LLM)生成的文本,近期有研究提出为博弈论场景中的智能体开发水印技术。然而,现有博弈智能体水印技术适用范围或能力有限,仅适用于完美信息博弈,无法应用于更丰富的博弈类型。本文提出一种新的博弈智能体水印方法,具备三个特性:a)可应用于不完美信息场景;b)直接集成到学习过程本身;c)仅产生有界的可利用性代价。为此,本文引入扰动遗憾最小化方法,在观察前对效用添加扰动,以促使学习算法嵌入水印。实验表明,该水印仅产生极小的可利用性代价,且在以人类速度进行仅几小时的博弈后即可被检测到。
英文摘要
Many real-world interactions among self-interested parties can be modeled by game theory, and the rapid advancements in AI have raised concerns about the possible misuse---accidental or deliberate---of superhuman or human-level game-playing agents by bad actors. While AI watermarking has mainly been applied to LLM-generated texts, a recent line of work proposes developing watermarking techniques for agents in game-theoretic settings. However, existing watermarking techniques for game-theoretic agents are not readily applicable due to their limited scope or capabilities---they are tailored to perfect-information games and are thus inapplicable to richer game types. We propose a new approach to watermarking game-playing agents, which a) can be applied to imperfect-information settings; b) is directly integrated into the learning process itself; and c) incurs only a bounded cost in exploitability. For this purpose, we introduce perturbed regret minimization, which adds perturbations to the utilities prior to observation so as to encourage the learning algorithm to embed the watermark. Our experiments show that the watermark incurs only a small exploitability cost and can be detected within just a couple of hours of gameplay at human speed.