arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

求解器引导的混合策略均衡推理

Solver-Guided Reasoning for Mixed-Equilibrium Strategies

Han Wang, Philippe Beardsell, Boning Li, Aaron Sasmita, Shuai Li, Hongyuan Zha, Baoxiang Wang

arXiv 2608.06741首次发表:更新:

发表机构

Shanghai Jiao Tong University; GTO Wizard; Tsinghua University; The Chinese University of Hong Kong, Shenzhen; Vector Institute(上海交通大学; GTO Wizard; 清华大学; 香港中文大学(深圳); 向量研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出混合策略决策树(MDT),用求解器输出引导LLM在博弈中的均衡推理,在无限制德州扑克8种LLM配置下,将与均衡的L1距离降低52.6%,且策略可移植性良好。

AI 中文摘要

大型语言模型(LLM)的推理通常基于人类文本、人类演示和人类生成的理由。然而,对于复杂博弈中的均衡推理,依赖人类数据可能并非最优。事实上,人类博弈往往受直觉和启发式规则引导,可能与博弈均衡存在显著偏差。在具有混合策略均衡的博弈中,这种偏差会被放大,因为人类数据严重偏向纯策略。因此,用此类数据训练LLM会产生较弱的博弈策略。为赋予LLM博弈推理能力,本研究探讨如何利用求解器输出引导均衡博弈。我们提出混合策略决策树(Mixed-Strategy Decision Tree, MDT),将均衡的隐性最优性转化为人类和LLM均可理解的稀疏策略规则。使用求解器输出而非人类标注,使我们能将输入扩展到任意新状态和延续。我们在无限制德州扑克(No-Limit Texas Hold'em, NLH)上实例化该研究,向求解器查询超过2.5亿个混合策略决策;MDT与其他技术结合,在8种不同LLM配置下,将与均衡的L1距离降低了52.6%。仅路由的消融测试评估了基于影子的对比的增量贡献,完整河牌终局和骰子游戏(Liar's Dice)实验则评估了策略保真度和在原始NLH通信设置之外的可移植性。

英文摘要

Reasoning in large language models (LLMs) is often grounded in human text, human demonstrations, and human-generated rationales. For equilibrium reasoning in complex games, however, relying on human data can be suboptimal. In fact, human play is often guided by intuition and heuristics and can deviate substantially from game equilibrium. This discrepancy is amplified in games with mixed-strategy equilibria, where human data is heavily biased toward pure strategies. Consequently, conditioning LLMs on this data yields weak game strategies. To grant LLMs the reasoning capacity in games, in this work, we study how to elicit equilibrium play using solver output. We propose Mixed-Strategy Decision Tree (MDT), which articulates the silent optimality of the equilibrium into sparse strategic rules that both humans and LLMs could understand. Using solver output rather than human annotation allows us to extend the input to arbitrarily new states and continuations. We instantiate this study on No-Limit Texas Hold'em by querying a solver oracle for over \textbf{250 million mixed-strategy decisions}; MDT together with other techniques \textbf{reduces the $\ell_1$ distance to the equilibrium by $52.6\%$} across $8$ different LLM configurations. A Route-only ablation tests the incremental contribution of the shadow-based contrast, while complete River-endgame and Liar's Dice experiments evaluate strategic fidelity and portability beyond the original NLH communication setting.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑