发表机构
Fudan University; Shanghai Key Laboratory of Data Science; College of Computer Science and Artificial Intelligence, Fudan University(复旦大学; 上海市数据科学重点实验室; 复旦大学计算机科学与人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长时程任务中上下文管理瓶颈,提出ContextEvo框架,从轨迹中学习上下文策略,在多个基准上超越或媲美主流智能体框架。
AI 中文摘要
框架进化通过从执行轨迹中学习来改进LLM智能体,但现有的基于经验和技能的方法在长时程任务上效果较差。随着交互的增多,有用的证据可能被冗余或过时的上下文所淹没,使得上下文管理本身成为关键瓶颈。我们提出了ContextEvo,一个从长时程轨迹中学习上下文策略的框架。ContextEvo在关键决策点重建模型可见的上下文,识别与上下文相关的失败,并应用有针对性的策略更新。基于开源的Pi-agent框架,ContextEvo在三个长时程任务基准上提升了性能,取得了与多个著名智能体框架(包括Codex、OpenCode和OpenClaw)相当或更好的结果。进一步的分析表明,固定的或局部进化的上下文策略在长时程信息压力下可能表现不佳,而我们的方法能适应每个环境的信息需求。
英文摘要
Harness evolution improves LLM agents by learning from execution trajectories, but existing experience- and skill-based methods are less effective on long-horizon tasks. As interactions grow, useful evidence can be buried by redundant or outdated context, making context management itself a key bottleneck. We introduce ContextEvo, a framework that learns a context policy from long-horizon trajectories. ContextEvo reconstructs the model-visible context at key decision points, identifies context-related failures, and applies targeted policy updates. Starting from the open-source Pi-agent harness, ContextEvo improves performance across three long-horizon task benchmarks, achieving results comparable to or better than several prominent agent harnesses, including Codex, OpenCode, and OpenClaw. Additional analyses show that fixed or locally evolved context strategies can fall short under long-horizon information pressure, while our methods adapt to the information demands of each environment.
Comments27 pages, 6 figures