arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17710cs.LG

使用Transformer进行规划:计算链与结构化上下文窗口

Planning with Transformers: Chain of Computation and Structured Context Windows

Ehsan Futuhi, Nathan R. Sturtevant

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型解决规划问题的不足,提出计算链架构,利用结构化上下文窗口,让语言模型学习规划策略、预测世界模型并执行算术运算,在多个任务上取得高成功率,能解决复杂汉诺塔问题且减少训练数据。

中文摘要 AI 辅助

大语言模型(LLMs)在机器学习诸多领域产生了显著影响,但近期研究表明它们在可靠解决规划问题上存在困难。同时,理论结果显示作为现代LLMs核心架构的Transformer是图灵完备的。本文研究了LLMs理论计算能力与其经验规划性能之间的明显差距。提出了计算链(COC),将基于Transformer的语言模型置于迭代循环中,利用其模式匹配系统的优势。COC使用结构化上下文窗口(SCW),在每个规划步骤支持选择使用哪个窗口。在该架构内,语言模型能学习规划策略、预测世界模型并执行规划所需的算术运算。研究表明,给定类似图灵机磁带的追加式SCW时,即使从头训练的相对小的语言模型也能学习规划策略并从少量训练实例中进行泛化,在BlocksWorld和煎饼谜题上成功率超过99.89%。对汉诺塔失败案例的分析表明,失败源于算术运算或遇到未见令牌。通过对算术的符号支持规划或为SCW使用确定性下推自动机(PDA)公式,COC能解决多达20个圆盘的汉诺塔问题实例,需要超过100万个动作,且所需训练数据大幅减少。

英文摘要

Large Language Models (LLMs) have had a remarkable impact across many areas of machine learning. However, recent studies have shown that they struggle to reliably solve planning problems. At the same time, theoretical results have shown that transformers, the core architecture underlying modern LLMs, are Turing-complete. In this work, we investigate this apparent gap between the theoretical computational power of LLMs and their empirical planning performance. We propose Chain of Computation (COC), a computational architecture that places a transformer-based LM inside an iterative loop, leveraging its strength as a pattern-matching system. The COC uses a Structured Context Window (SCW) which provides a constant-sized context window with support for choosing which window is used at each planning step. Within this architecture, the LM is able to learn a planning policy, predicts the world model, and performs the arithmetic operations required during planning. We show that, when given an append-only SCW (resembling a Turing Machine tape), even relatively small LMs trained from scratch can learn planning policies and generalize from a small number of training instances within each planning domain, achieving success rates above 99.89\% on BlocksWorld and the Pancake puzzle. Our analysis of failure cases in Tower of Hanoi (TOH) reveals that they arise from arithmetic operations or from encountering previously unseen tokens. We show that COC can solve TOH problem instances with up to 20 disks, requiring over 1 million actions, while requiring substantially less training data by either (1) planning with symbolical support for arithmetic or by (2) using a deterministic pushdown automaton (PDA) formulation for the SCW.

发表机构

  • Department of Computing Science, University of Alberta(阿尔伯塔大学计算科学系)
  • Alberta Machine Intelligence Institute (Amii)(阿尔伯塔机器智能研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑