发表机构
Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出战略延续优化(SCO)策略构建方法,通过锦标赛实验对比发现,SCO 策略比基于解析 ICM 的策略更优,且在替换对手为 LLMs 等时仍保持优势,明确了 ICM 作为锦标赛策略构建目标的局限性。
AI 中文摘要
独立筹码模型(ICM)将锦标赛筹码转换为参考奖金权益,策略通常基于这些值构建。由于 ICM 仅读取筹码量,它忽略了行动顺序、盲注义务和座位轮换,且未对大筹码量玩家可能淘汰短筹码玩家所施加的淘汰压力进行定价。这些遗漏会改变决定行动的后继状态对比。我们提出战略延续优化(SCO),这是一种策略构建方法,它枚举当前手牌结果,将其映射到后继状态,并用有限锦标赛模型计算的延续值对这些状态定价,进而优化并冻结最终的当前手牌策略。固定 ICM 对比策略仅改变一点:相同的优化器用解析 ICM 对后继状态定价来求解同一游戏,因此两种策略的差异仅源于定价方式。我们在奖池为 100 万美元的三人全押/弃牌锦标赛中评估所得策略。相对于冻结的战略延续基准,解析 ICM 在全部 2838 个状态-座位条目中的平均绝对值误差为 9036 美元。该价值误差会重写其定价范围:相对于每个决策点自身的固定 ICM 全押范围,SCO 的全押频率平均变动 14.08%。为对这些不同行动定价,我们在改变 focal 策略、固定对手和延续评估器的同时,对比全部 946 个状态和 3 个策略持有者。SCO 生成的策略平均每手牌多获得 214.33 美元奖金权益,且在全部 2838 个匹配单元中占优 2433 个。即使将求解器构建的对手替换为两个大语言模型(LLMs)和一组非建模阈值玩家,该排序依然成立。这一价值-策略-成本链直接表明,ICM 何时会成为锦标赛策略构建的不充分目标。
英文摘要
The Independent Chip Model (ICM) converts tournament chips into reference prize equity, and policies are routinely constructed against those values. Because ICM reads only stack sizes, it omits action order, blind obligations, and seat rotation, and it does not price the elimination pressure a big stack puts on the short stacks it can bust. Those omissions can alter the successor-state contrasts that determine a move. We introduce Strategic-Continuation Optimization (SCO), a policy-construction method that enumerates current-hand outcomes, maps them to successor states, prices those states with continuation values computed from the finite tournament model, and optimizes and freezes the resulting current-hand policy. The fixed-ICM comparison policy changes one thing only: the same optimizer solves the same game with successor states priced by analytic ICM, so the two policies differ only through that pricing. We evaluate the resulting policies in a three-player jam/fold tournament with a \$1M prize pool. Relative to the frozen strategic-continuation benchmark, analytic ICM has \$9{,}036 mean absolute value error across all 2,838 state--seat entries. That value error rewrites the ranges it prices: measured against each decision point's own fixed-ICM jam range, SCO moves the jam frequency by an average of 14.08\%. To price those different moves, we compare all 946 states and three policy owners while changing only the focal policy and holding both opponents and the continuation evaluator fixed. The policy produced by SCO earns \$214.33 more prize equity per hand on average and is favored in 2,433 of 2,838 matched units. The ordering survives replacing the solver-built opponent with two LLMs and with a family of non-modeling threshold players. This value-to-policy-to-cost chain shows directly when ICM becomes an inadequate objective for tournament strategy construction.
Comments34 pages, 2 figures