arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoupVisor:基于回合与挑战决策支持的策略优化

CoupVisor: Strategy Optimization by Round and Challenge Decision Support

Cris Huynh

arXiv 2608.15868首次发表:更新:

AI 中文总结

本文提出用于卡牌游戏《Coup》的决策支持系统CoupVisor,其结合角色可能性与持牌数估计声明真实性,实验发现以获胜为导向的奖励策略优于基准方法。

AI 中文摘要

本文提出了CoupVisor,这是一款用于隐藏信息卡牌游戏《Coup》的决策支持系统。它解决两个问题:玩家在每回合应采取何种行动,以及何时对对手的声明提出挑战。该系统围绕对游戏事件的单一描述构建,该描述可在手动游戏、已记录游戏的重放、模拟、信念跟踪、顾问建议以及基于学习的策略之间共享。CoupVisor通过结合每个角色的可能性以及声明者仍持有的卡牌数量来估计声明为真的概率,这纠正了游戏中首个声明被标记为可疑却无任何证据的情况。我们在大量模拟游戏和不同对手风格下,将遵循规则的顾问与若干学习型及启发式玩家进行比较。主要发现是,奖励的选择——无论是奖励短期收益还是最终赢得游戏——决定了哪种学习方法表现最佳,且以获胜为导向的奖励产生的策略优于所有基准方法。

英文摘要

This paper presents CoupVisor, a decision-support system for the hidden-information card game Coup. It addresses two questions: what a player should do on each turn, and when a player should challenge an opponent's claim. The system is built around a single description of game events, which is shared across manual play, replay of recorded games, simulation, belief tracking, advisor recommendations, and learning-based policies. CoupVisor estimates the chance that a claim is truthful by combining how likely each role is with how many cards the claimant still holds, which corrects a case where the very first claim of a game was flagged as suspicious despite no evidence. We compare a rule-following advisor and several learned and heuristic players across many simulated games and different opponent styles. Our main finding is that the choice of reward, whether it rewards short-term gains or ultimately winning the game, decides which learning approach performs best, and that a win-oriented reward produces a policy that outperforms all baselines.

Comments15 Pages and 9 pages of appendix

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑