arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于经验记忆提升大语言模型智能体的序列决策能力

Towards Improving Sequential Decision-Making in LLM Agents via Experience Memory

Jakub Rada, Viliam Lisý

arXiv 2608.03420首次发表:更新:

AI 中文总结

本研究针对LLM智能体序列决策性能不足的问题,提出带经验记忆的智能体框架,通过对局后反思与规则提取,在不修改模型权重的情况下提升了井字棋任务的表现。

AI 中文摘要

大语言模型(LLM)在单次推理任务上已取得显著进步,但其序列决策表现仍有待深入研究。我们在完全可观测的双人零和博弈场景中开展研究,该场景可提供真实评估:博弈结果由规则决定,单个动作的最优性可被计算或近似,无需依赖评判模型。在各模型层级中,LLM在井字棋、四子棋等简单博弈中表现次优,会输给基于蒙特卡洛树搜索(MCTS)的对手。在保留博弈树结构但改写其表层形式的混淆设置下,模型性能基本未变,表明性能差距并非完全由记忆策略的召回能力导致。受此性能差距驱动,我们提出一种增强了经验记忆的智能体框架,该框架专为序列场景设计,用于解决信用分配等序列决策的常见挑战。我们发现,无需修改模型权重,仅通过对局后反思与规则提取即可在井字棋任务上取得可测量的性能提升。

英文摘要

Large language models have improved substantially on single-shot reasoning tasks, but their performance in sequential decision-making is less well understood. We study this on fully-observable two-player zero-sum games, which provide ground-truth evaluation: outcomes are determined by the rules, and optimality of individual moves can be computed or approximated, without relying on a judge model. Across model tiers, LLMs play suboptimally in simple games such as tic-tac-toe or Connect Four, and lose to MCTS opponents. Obfuscations that preserve the game tree but rewrite its surface form leave performance largely unchanged, indicating the gap is not fully explained by recall of memorized strategies. Motivated by this performance gap, we introduce an agentic framework enhanced with an experience memory designed for the sequential setting and addressing common challenges of sequential decision-making such as credit assignment. We show that post-game reflection and rule extraction yield measurable improvements on tic-tac-toe without modifying the model weights.

Comments8 pages, 6 figures, 14 tables, 5 appendices, accepted at Neuro-Symbolic Intelligence for LLMs and Autonomous Agents workshop at IJCAI 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑