长时程智能体架构:层级、时钟与级联智能
An Architecture for Long-Horizon Agents: Levels, Ticks and Cascaded Intelligence
浏览论文内容
中文总结 AI 辅助
针对语言模型智能体难以执行跨天或数周长时程任务的问题,本文提出由按时间尺度分层的层级、时钟驱动的tick和级联智能组成的三部分架构,通过十天实验证明该架构能使智能体在人类每日关注一次的情况下复现强化学习结果,并实现持续运行与知识积累。
中文摘要 AI 辅助
语言模型智能体正越来越多地被要求执行跨越数天或数周的工作,例如运维修复或研究项目。此类任务超出了任何上下文窗口、任何进程以及任何人能够关注的时间间隔。在本文中,我们认为,长时程智能体必须在不遗忘的情况下持续运行,然后才能进行持续学习。这种能力存在于模型周围的框架(harness)中,而非模型本身。我们从长时程场景中推导出七个瓶颈,并用一个由三部分组成的层次架构来应对:(i)按时间尺度索引的层级,每层维护一个有限的文件,总结其下一层的内容;(ii)一个时钟驱动的时钟(tick)作为自主行动的单位;(iii)级联智能,即工作仅在审查失败后才升级到能力更强的模型。我们报告了一项为期十天的活动,在该活动中,基于此架构构建的智能体在人类每天关注一次的情况下复现了一个已发表的强化学习结果,并展示了:(1)该智能体在活动的每次上下文重置和会话边界中保持了线索;(2)早期写入的操作知识在不改变模型权重的情况下改变了后续行为;(3)学习组件将进入此类系统的位置。总体而言,我们的经验表明,这些智能体的持续学习需要一个超越每个上下文和进程的底层支撑,而框架已运行的检查正是学习器所属之处。
英文摘要
Language-model agents are increasingly asked to carry out work spanning days or weeks, such as an operations remediation or a research programme. Such a task outlives any context window, any process and any interval at which a person can attend. In this paper, we argue that a long-horizon agent must run continually without forgetting before it can learn continually. This ability lies in the harness around the model rather than in the model itself. We derive seven bottlenecks from the long-horizon setting and answer them with a hierarchical architecture of three parts: (i) levels indexed by time scale, each keeping a bounded file summarising the level below; (ii) a clocked tick as the unit of autonomous action; and (iii) cascaded intelligence, where work is escalated to a more capable model only after failing review. We report on a ten-day campaign in which an agent built on this architecture reproduced a published reinforcement-learning result with a human attending once a day, and show (1) the agent kept the thread across every context reset and session boundary of the campaign, (2) operating knowledge written early changed later behaviour with no change to model weights, and (3) where learned components would enter such a system. Overall, our experience suggests continual learning for these agents needs a substrate outliving every context and process, and the checks the harness already runs are where a learner belongs.
发表机构
- Salesforce AI Research(赛富时人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。