arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAGE:用于高效LLM游戏对弈的结构化策略推理

SAGE: Structured Strategic Reasoning for Efficient LLM Game Playing

Zhiwei Chen, Tianchun Wang, Zhongtao Rao, Haiming Zhu, Ding Cao, Tianxiang Zhao

arXiv 2609.34342首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou); Johns Hopkins University; Microsoft; Fudan University; University of Science and Technology of China(香港科技大学(广州); 约翰斯·霍普金斯大学; 微软; 复旦大学; 中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SAGE是一个无需训练的推理时框架,通过锚定、适应和重新校准三个操作结构化LLM策略推理,在多个不完美信息博弈中显著提升收益并减少令牌使用。

AI 中文摘要

一个强大的LLM策略智能体应当能够对不确定的未来进行前瞻性推理,根据对手的行为倾向调整自身策略,并从交互经验中持续重新校准其决策过程。然而,将这些来源纳入自由形式的推理可能导致无根据的策略假设、不一致的对手估计,以及来自无关历史交互的有害干扰。为解决这些问题,我们提出了SAGE,一种无需训练的推理时框架,该框架围绕三个协调操作来结构化LLM策略推理:锚定(anchor)、适应(adapt)和重新校准(recalibrate)。SAGE首先将推理锚定在均衡策略上,该策略提供了具有策略有效性的先验。然后,它基于对手行为倾向的软信念来调节偏离该锚定的偏差,从而实现针对对手的利用。最后,SAGE将策略相关的交互提炼为关于先前缺失考虑的假设性反事实,使过去的经验能够重新校准模型的推理。我们在三个重复的不完美信息博弈中评估了SAGE:Leduc Hold'em、Liar's Dice和Goofspiel,针对每种博弈中的多种对手类型。与推理密集型的LLM智能体(包括Suspicion-Agent、ReTA、Agent-Pro、EMO和Hypothetical Minds)相比,SAGE在Liar's Dice中实现了高达127.6%的收益提升,同时分别将输入和输出令牌使用量减少了高达80%和90%。在直接对局中,SAGE在Leduc Hold'em、Liar's Dice和Goofspiel中分别对5/10、8/10和8/10的评估对手实现了非负的平均收益,同时使用了相对较少的令牌。代码可在该https URL获取。

英文摘要

A strong LLM strategic agent should reason prospectively over uncertain futures, adapt its strategy to opponents' behavioral tendencies, and continuously recalibrate its decision process from interaction experience. However, incorporating these sources in free-form reasoning could lead to unsupported strategic assumptions, inconsistent opponent estimates, and harmful interference from irrelevant historical interactions. To address these issues, we propose SAGE, a training-free inference-time framework that structures LLM strategic reasoning around three coordinated operations: anchor, adapt, and recalibrate. SAGE first anchors reasoning to an equilibrium policy that provides a strategically valid prior. It then conditions deviations from this anchor on a soft belief over opponent behavioral tendencies, enabling opponent-specific exploitation. Finally, SAGE distills strategically related interactions into counterfactual hypotheses about previously missing considerations, allowing past experience to recalibrate the model's reasoning. We evaluate SAGE on three repeated imperfect-information games: Leduc Hold'em, Liar's Dice, and Goofspiel, against various opponent types in each game. Compared with reasoning-intensive LLM agents, including Suspicion-Agent, ReTA, Agent-Pro, EMO, and Hypothetical Minds, SAGE achieves up to a 127.6% payoff improvement in Liar's Dice while reducing input and output token usage by up to 80% and 90%, respectively. In direct match-up play, it attains non-negative mean payoff against 5/10, 8/10, and 8/10 evaluated opponents in Leduc Hold'em, Liar's Dice, and Goofspiel, respectively, while using relatively fewer tokens. Code is available at https://github.com/chenzhwsysu57/SAGE.

Comments31 pages with multiple figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑