无状态语言智能体:扩展长时程自动化研究
Stateless Language Agents: Scaling Long-Horizon Automated Research
浏览论文内容
中文总结 AI 辅助
针对长时程自动化研究中智能体状态累积导致的低效问题,提出无状态语言智能体框架,通过有状态搜索与无状态智能体分离,在十亿令牌预算下以更少资源超越现有方法。
中文摘要 AI 辅助
自动化研究系统日益在长时间跨度上运行大语言模型智能体,但更多的推理本身并不会带来更多的进展:智能体会重放不断增长的历史记录,重复彼此的工作,或在令牌消耗持续进行时停止实验。然而,大多数评估使用短预算或早期饱和的基准,使得这些失败模式未被测试。我们将这些失败追溯到两个选择:研究状态存储在哪里,以及由谁决定下一步尝试什么。我们引入了无状态语言智能体(SLAs),其构建原则是“有状态搜索,无状态智能体”:没有任何智能体会跨调用携带其对话;相反,执行框架拥有研究状态(候选解决方案和测量结果),并为每次调用重建一个全新的、角色特定的上下文。每个智能体所看到的内容成为一种显式设计选择,而非随运行而增长的历史记录。我们在SLA框架中实现了这一原则,其中无状态的顾问读取跨搜索方向的框架总结证据,并将具体实验分配给并行的工作者。我们在软件工程、内核优化和算法设计任务上,以高达十亿令牌的预算,将SLA与三个近期框架进行了评估。SLA在每项任务上都取得了最佳最终结果,并以少于84%的令牌达到了最强内核基线的最终性能。来自共享检查点的消融实验表明,聚焦上下文和显式分配各自对SLA的进展有所贡献,其效果在完整运行中可能复合,而顾问消耗的令牌不到0.6%。这些结果支持SLA,即将持久的研究状态保留在智能体对话之外,并表明短评估周期可能误判研究系统及其组件。
英文摘要
Automated research systems increasingly run LLM agents over long horizons, but more inference does not by itself produce more progress: agents replay growing histories, duplicate one another's work, or stop experimenting while token consumption continues. Yet most evaluations use short budgets or benchmarks that saturate early, leaving these failure modes untested. We trace these failures to two choices: where research state lives and who decides what to try next. We introduce Stateless Language Agents (SLAs), built on the principle of stateful search with stateless agents: no agent carries its conversation across invocations; instead, the harness owns the research state (candidate solutions and measured outcomes) and reconstructs a fresh and role-specific context for every invocation. What each agent sees becomes an explicit design choice rather than a history that grows with the run. We implement this principle in the SLA framework, where a stateless Advisor reads harness-summarized evidence across search directions and assigns concrete experiments to parallel Workers. We evaluate SLA against three recent frameworks on software engineering, kernel optimization, and algorithm design at budgets of up to one billion tokens. SLA achieves the best final result on every task and reaches the strongest kernel baseline's final performance with over 84% fewer tokens. Ablations from shared checkpoints show that focused contexts and explicit assignments each contribute to SLA's progress, with effects that can compound over full runs, while the Advisor consumes less than 0.6% of tokens. These results argue for SLAs, which keep durable research state out of agent conversations, and show that short evaluation horizons can misjudge research systems and their components.
发表机构
- Stanford University(斯坦福大学)
- Carnegie Mellon University(卡内基梅隆大学)
- University of Washington(华盛顿大学)
- SambaNova Systems, Inc.(SambaNova系统公司)
- UC Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。