发表机构
Duke University(杜克大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长程任务中LLM智能体的挑战,提出PRO-LONG框架,通过程序化内存管理上下文,保留交互日志并利用编码智能体进展高效搜索历史,在ARC-AGI-3上有显著提升,减少令牌使用并达到高准确率。
AI 中文摘要
长程任务对大语言模型(LLM)智能体构成持续挑战,在如ARC-AGI-3等连续学习基准测试中表现受限。现有多种智能体框架来弥补差距,但在上下文管理上面临权衡。本文提出PRO-LONG,一个围绕程序化内存构建的最小化上下文管理框架。它通过保留完整结构化交互日志,并利用编码智能体的进展高效搜索历史记录来解决权衡问题。在ARC-AGI-3公共游戏集上,PRO-LONG在前沿模型上比基础编码智能体平均提高18.0个百分点,在使用更少令牌的情况下匹配或超越最先进的专用框架。与Fable 5配合时,PRO-LONG以1750美元的总成本实现了97.4%的最佳@2。相关代码和日志可在指定网址获取。
英文摘要
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents. This gap is reflected in their limited performance on continual learning benchmarks such as ARC-AGI-3, especially when models are evaluated out of the box. Various agent harnesses have been proposed to close this gap, and each commits to a strategy for handling long sequences of observations, i.e., what information to save from the environment and how to load it into model context, a choice we argue is particularly consequential. Existing methods for context management face a significant tradeoff, as preserving more information makes retrieving relevant details less tractable. We propose PRO-LONG, a minimal context management framework built around programmatic memory for LLM agents in long-horizon, exploratory settings. PRO-LONG addresses the tradeoff by keeping a complete, structured interaction log and capitalizing on recent progress in coding agents to search this history efficiently. On the full ARC-AGI-3 public game set, PRO-LONG improves over a base coding agent by an average of 18.0 percentage points across frontier models, and matches or exceeds state-of-the-art specialized harnesses (up to 76.1% pass@1) while using 4.2-5.8x fewer tokens. With Fable 5, PRO-LONG achieves 97.4% best@2 at a total cost of \$1,750. Relevant code and logs are available at https://github.com/alexisfox7/PRO-LONG.