大型语言模型(LLMs)并非优秀的战略家,然而记忆增强的智能体可提升推理能力
LLMs Are Not Good Strategists, Yet Memory-Enhanced Agency Boosts Reasoning
浏览论文内容
中文总结 AI 辅助
该研究针对LLM在长时环境中战略推理的缺陷,提出EpicStar框架,结合跨回合记忆与动态门控机制,在星际争霸II测试中实现更高胜率且令牌消耗大幅减少。
中文摘要 AI 辅助
大型语言模型(LLMs)在长时环境中的战略推理往往受限于子目标不一致的问题。在这类场景中,有限的注意力资源会阻碍模型在数千步的推理过程中保持战略连贯性,这一限制会导致战略漂移,即局部决策无法在整个推理过程中维持连贯的轨迹。为解决该问题,我们提出EpicStar框架,该框架使智能体能够将记忆作为策略进行学习,以应对长时推理任务。具体而言,智能体维护一个成功过往回合的库作为启发式工具,同时通过工作记忆跟踪短期环境变化;推理阶段,动态门控机制决定是直接执行检索到的动作,还是通过将检索到的回合与当前工作记忆进行上下文融合来开展新的推理。我们以星际争霸II(StarCraft II)为测试平台,将EpicStar与多种对手风格进行评估,结果显示其显著优于基线方法,在实现更高胜率的同时,消耗的令牌数减少了一个数量级,且在不同难度级别和对手策略下均能持续保持这一优势。我们的研究结果有力证明,结构化的跨回合记忆对于使LLM智能体在动态自主环境中执行稳健的长期战略至关重要。
英文摘要
Strategic reasoning in Large Language Models (LLMs) within long-horizon environments is often limited by inconsistent subgoals. In these settings, finite attention resources prevent the model from maintaining strategic coherence over thousands of steps. This limitation leads to strategic drift, where localized decisions fail to sustain a coherent trajectory across reasoning. To address this, we introduce EpicStar, a framework that enables agents to learn memory as policy to tackle long-horizon reasoning. Specifically, the agent maintains a bank of successful past episodes as a heuristic alongside a working memory to track short-term environmental changes. During inference, a dynamic gating mechanism determines whether to execute a retrieved action directly or to perform new reasoning through a contextual fusion of the retrieved episodes and current working memory. Utilizing StarCraft II as the testbed, we evaluated EpicStar against diverse opponent styles. It significantly outperforms baseline methods, achieving higher win rates while consuming an order of magnitude fewer tokens, and it maintains this advantage consistently across difficulty levels and opponent strategies. Our findings provide compelling evidence that structured cross-episode memory is essential for enabling LLM agents to perform robust, long-term strategic execution in dynamic, autonomous settings.
发表机构
- University of Chicago(芝加哥大学)
- University of Wisconsin-Madison(威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。