发表机构
Georgia Institute of Technology; Etude AI(佐治亚理工学院; Etude人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出语言模型智能体机器(LAM)抽象,量化LLM智能体支架的计算成本,通过通信、访问、重计算和可靠性四类结果建立资源理论,并在GPT-6 Astra上验证预测。
AI 中文摘要
语言模型智能体日益依赖于管理有界上下文、持久记忆、工具、验证和重复执行的“支架”(harness),然而现有的模型能力概念并未量化这些机制所消耗的计算资源。我们引入了语言模型智能体机器(Language Model Agent Machine, LAM),这是一种资源有界的抽象,它固定了底层语义模型,同时明确对支架级资源收费。我们建立了四类结果。通信:在同时调用-传输预算下,LAM执行在实例级别上等价于红-蓝鹅卵石(red-blue pebbling)问题,将经典的I/O下界转移到上下文-内存流量上。访问:内存接口导致渐近分离,包括在指针追逐上随机访问与非推测性顺序访问之间的Θ(n)差距。重计算:位反转DAG在上下文容量C和持久内存容量S下需要Θ(n^2/(C+S)+n)次模型调用,量化了存储中间状态何时避免重复的语义计算。可靠性:我们推导出严格的阶段局部采样界限、精确的不完美验证成本,以及一个Young-Daly型检查点定律,具有闭式最优验证间隔。在GPT-6 Astra上的受控和留出实验测试了通信和可靠性预测,包括检查点最优性、程序化检查下的策略选择,以及在链式MATH任务中调用粒度、逻辑输入流量和可靠性之间的权衡。总之,这些结果为语言模型智能体支架的计算成本提供了资源理论。
英文摘要
Language-model agents increasingly rely on harnesses that manage bounded context, persistent memory, tools, verification, and repeated execution, yet existing notions of model capability do not quantify the computational resources these mechanisms consume. We introduce the Language Model Agent Machine (LAM), a resource-bounded abstraction that fixes the underlying semantic model while explicitly charging harness-level resources. We establish four classes of results. Communication: LAM execution is instancewise equivalent to red--blue pebbling under simultaneous call--transfer budgets, transferring classical I/O lower bounds to context--memory traffic. Access: memory interfaces induce asymptotic separations, including a $Θ(n)$ gap between random and non-speculative sequential access on pointer chasing. Recomputation: bit-reversal DAGs require $Θ(n^2/(C+S)+n)$ model calls with context capacity $C$ and persistent-memory capacity $S$, quantifying when stored intermediate state avoids repeated semantic computation. Reliability: we derive tight stage-local sampling bounds, exact imperfect-verification costs, and a Young--Daly-type checkpoint law with a closed-form optimal verification interval. Controlled and held-out experiments on GPT-6 Astra test communication and reliability predictions, including checkpoint optima, policy selection under programmatic checking, and tradeoffs among call granularity, logical input traffic, and reliability on chained MATH tasks. Together, these results provide a resource theory for the computational cost of language-model agent harnesses.