arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06816cs.AI

在《魔法与代码传奇》中使用策略与价值网络进行非健全搜索

Unsound Search with Policy and Value Networks in Legends of Code and Magic

Dustin Rubin

首次发表
浏览论文内容

中文总结 AI 辅助

本文在《魔法与代码传奇》中提出基于策略与价值网络的非健全搜索方法,以51.35%胜率击败冠军ByteRL,并增强了对抗攻击的韧性。

中文摘要 AI 辅助

在具有可枚举信念状态的完美信息和不完美信息游戏中,决策时搜索是游戏AI的有效方法。集换式卡牌游戏是具有大信念状态的不完美信息游戏。《魔法与代码传奇》是一场集换式卡牌游戏竞赛,其信念状态为$2^{101}$。《魔法与代码传奇》(LoCM)冠军ByteRL在游戏中不使用搜索。其他研究声称,由于信念状态的数量,基于枚举的健全搜索在该类型游戏中不可用。我们测量了先前定义的三个属性,这些属性预测了理论上非健全的完美信息蒙特卡洛缺陷代价较低的情况,发现LoCM处于有利区域。从亚军策略NeteaseOPD的模仿学习开始,我们创建了一个策略和价值前馈网络。我们的智能体在基于亚军草稿构建的对手牌组先验上,对从该先验采样的世界进行搜索。在战斗阶段使用我们最严格的配置,我们在10,000场预注册比赛中,使用LoCM官方裁判和时间限制,以51.35%的胜率(95%置信区间[50.37, 52.33])击败了ByteRL。搜索在我们智能体与ByteRL的对局中并非次要因素。没有搜索时,该智能体得分26.8%,添加搜索后增加了+24.6个百分点。在不完美信息游戏中进行非健全搜索可能是可利用的。我们复现了针对ByteRL的已发表的最佳响应攻击。然后我们将相同的攻击协议应用于我们智能体的两种搜索配置,每种配置在每次迭代中都比ByteRL更好地抵抗攻击。在LoCM中,非健全搜索使我们获得了更强大、更具韧性的智能体。

英文摘要

Decision-time search in perfect and imperfect information games with enumerable belief states are effective methods for game AI. Collectible card games are imperfect information games with large belief states. Legends of Code and Magic is a collectible card game competition where the belief states are $2^{101}$. The Legends of Code and Magic (LoCM) champion, ByteRL, plays with no search. Other works claim sound enumeration-based search is unusable in the genre due to the number of belief states. We measured three previously defined properties that predict where theoretically unsound perfect information Monte Carlo's defects are cheap and found LoCM sits in the favorable region. Starting with imitation learning of the runner-up policy, NeteaseOPD, we created a policy and value feed-forward network. Our agent searches over worlds sampled from a prior over the opponent's deck built from the runner-up's drafts. Using our strictest configuration in the battle phase we beat ByteRL with a win percentage of 51.35% 95% CI [50.37, 52.33], over 10,000 pre-registered games using the LoCM official referee and time limit. Search is not a minor factor on the matchup between our agent and ByteRL. Without search this agent scores 26.8% and adding search adds +24.6 points. Unsound search in imperfect information games could be exploitable. We replicate a published best-response attack against ByteRL. We then apply the same attack protocol to two search configurations of our agent, and each one resists it better than ByteRL at every iteration. In LoCM unsound search gives us a stronger and more resilient agent.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑