抽象智能体
Abstraction Agent
- Institute for Interdisciplinary Information Sciences, Tsinghua University(清华大学交叉信息研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出零样本流程Abstraction Agent,利用LLM从自然语言博弈描述中发现特征并聚类,在HUNL等博弈上降低可利用性、击败基准,还可跨博弈转移,将LLM隐式知识转为下游算法的显式特征。
AI中文摘要:
信息抽象是将策略相似的私有状态分组为可处理数量的桶,对于将博弈求解算法扩展到大型不完美信息博弈至关重要。然而,构建有效的抽象传统上需要领域特定的评估器,如手牌强度计算器或胜率估计器,这需要专业知识和工程努力,且在大多数研究较少的博弈中不可用。我们提出了Abstraction Agent,一种零样本流程,它使用大型语言模型(LLM)从自然语言博弈描述中发现连续策略特征,对私有状态在这些特征上进行评分,并将其聚类为抽象桶,且在抽象构建期间无需任何特定博弈的评估器、训练数据或博弈树遍历。该流程分为四个阶段:带校准锚点的特征发现、批量私有状态评分、基于相关性的特征选择和k均值聚类。在 heads-up 无限注德州扑克(HUNL)转牌残局上,所得抽象将 lifted-strategy 可利用性相对于预期手牌强度基准降低了多达62%,并且在ROVER Trials(一个未出现在任何预训练语料库中的原创博弈)上,在每个粒度都击败了标量排名基准。除了这些定量基准外,该流程使用相同的提示转移到四注限注奥马哈、HUNL 翻前和翻后以及立直麻将,其中发现的特征跟踪每个博弈公认的策略概念。这是结构化知识 elicitation:将LLM参数中的隐式策略知识转换为显式数值特征,用于下游算法计算。代码可在该https URL获取。
英文摘要:
Information abstraction, which groups strategically similar private states into a tractable number of buckets, is essential for scaling game-solving algorithms to large imperfect-information games. Constructing effective abstractions, however, has traditionally required domain-specific evaluators such as hand-strength calculators or equity estimators, which demand expert knowledge and engineering effort and are unavailable for most less-studied games. We propose the Abstraction Agent, a zero-shot pipeline that uses a large language model (LLM) to discover continuous strategic features from a natural-language game description, score private states on these features, and cluster them into abstraction buckets, without any game-specific evaluator, training data, or game-tree traversal during abstraction construction. The pipeline runs in four phases: feature discovery with calibration anchors, batched private-state scoring, correlation-based feature selection, and $k$-means clustering. The resulting abstractions reduce lifted-strategy exploitability by up to 62% relative to an expected-hand-strength baseline on heads-up no-limit Texas hold'em (HUNL) turn endgames, and beat a scalar rank baseline at every granularity on ROVER Trials, an original game absent from any pretraining corpus. Beyond these quantitative benchmarks, the pipeline transfers with unchanged prompts to four-card Pot-Limit Omaha, HUNL preflop and flop, and Riichi Mahjong, where the discovered features track each game's recognized strategic concepts. This is structured knowledge elicitation: converting implicit strategic knowledge in LLM parameters into explicit numerical features for downstream algorithmic computation. The code is available at https://github.com/lbn187/AbstractionAgent.