arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05422cs.LGcs.MA

IFlowNets:将生成采样器扩展至不完美信息博弈中学习策略

IFlowNets: Extending Generative Samplers to Learn Strategies in Incomplete Information Games

Conor M. Artman, Nicholas Di, Scott Perkins

AI总结:

该研究将AFlowNets扩展为适用于不完美信息博弈的IFlowNets,解决了原有约束无法得到有效密度与训练目标的问题,在标准博弈环境中性能与速度优于或相当于OSMCCFR等方法。

AI中文摘要:

尽管许多算法将强化学习(RL)与反事实遗憾(CFR)方法相结合,以权衡计算速度与性能,但针对不完美信息博弈中博弈论应用的生成采样框架的研究较少。我们将生成流网络框架——对抗流网络(AFlowNets)扩展至不完美信息博弈,命名为信息流网络(IFNs)。我们证明,此前针对完美信息博弈中生成流网络确立的约束,无法得到对应玩家策略的有效密度及有效训练目标。我们表明,所提出的泛化方法IFlowNets缓解了该问题,且严格泛化了AFlowNets。在三个标准博弈环境的初步结果中,IFlowNets在性能与速度上与结果采样蒙特卡洛反事实遗憾(OSMCCFR)及标准RL方法表现相当或更优。

英文摘要:

While many algorithms blend reinforcement learning (RL) with counterfactual regret (CFR) methods to leverage tradeoffs in computational speed and performance, there are fewer investigations into generative sampling frameworks in game theoretic applications in incomplete information games. We extend a generative flow network framework, Adversarial Flow Networks (AFlowNets), to incomplete information games, called Information Flow Networks (IFNs). We prove that previously established constraints for generative flow networks in complete information games are inadmissible for obtaining valid densities (corresponding to player strategies) and a valid training objective. We show that our proposed generalization, IFlowNets, alleviates this issue and strictly generalizes AFlowNets. In preliminary results for three standard game environments, IFlowNets perform comparably to or better than Outcome Sampling Monte Carlo Counterfactual Regret (OSMCCFR) and standard RL-based methods in performance and speed.

补充信息

↑