arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ContextPilot:通过细粒度强化学习为智能体教授主动上下文管理

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

Zhuoshi Pan, Qizhi Pei, Junru Lu, Honglin Lin, H. Vicky Zhao, Di Yin, Xing Sun

arXiv 2608.28476首次发表:更新:

发表机构

Tsinghua University; Tencent; Shanghai AI Lab(清华大学; 腾讯; 上海人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对现有主动上下文管理方法的局限,提出ContextPilot框架,通过扩充工具集并设计细粒度RL方法,在长程智能体任务中实现更紧凑上下文下的更优性能,优于现有基线。

AI 中文摘要

长程智能体任务要求大语言模型(LLM)在多轮交互中迭代检索、整合并维护分散的信息,但保留所有交互历史会导致工作上下文持续增长。近期的主动上下文管理方法允许模型使用专用工具编辑自身工作上下文,但仍面临三个关键局限:(1)工具集受限,仅支持搜索、删除和摘要,不支持全局规划、长时记忆和自适应压缩;(2)探索效率低,对上下文管理动作统一处理,未考虑其对最终结果的异质性影响;(3)信用分配粒度粗,在强化学习(RL)中将最终轨迹级奖励分配给所有中间上下文编辑动作。为弥合这些差距,我们提出ContextPilot,一种用于长程智能体推理的主动上下文管理框架。我们的方法系统地为工具集增加规划、长时记忆和软上下文卸载工具。我们还提出一种针对上下文管理的RL方法,该方法利用上下文和熵变化识别关键编辑决策以进行分支采样,并从经过对应上下文编辑动作的所有分支轨迹中估计动作级优势。在长上下文问答和深度搜索任务上的实验表明,ContextPilot在工作上下文更紧凑的情况下实现了更强的性能,在各种基础模型和基准测试中始终优于现有基线。代码可在该https URL获取。

英文摘要

Long-horizon agentic tasks require large language models (LLMs) to iteratively retrieve, integrate, and maintain dispersed information across multi-turn interactions, but preserving all interaction histories leads to a continuously growing working context. Recent proactive context management methods allow models to edit their own working context with specialized tools, yet they still face three key limitations: (1) a limited toolset restricted to search, deletion, and summarization, with no support for global planning, long-term memory, and adaptive compression; (2) inefficient exploration that treats context management actions uniformly despite their heterogeneous impacts on final outcomes; and (3) coarse-grained credit assignment that assigns the final trajectory-level reward to all intermediate context editing actions during RL. To bridge these gaps, we introduce ContextPilot, a proactive context management framework for long-horizon agentic reasoning. Our approach systematically augments the toolset with planning, long-term memory, and soft context offloading tools. We further propose an RL method tailored for context management, which uses context and entropy variation to identify critical editing decisions for branch sampling and estimates action-level advantages from all branched trajectories that pass through the corresponding context editing action. Experiments on long-context QA and deep search tasks show that ContextPilot achieves stronger performance with a more compact working context, consistently outperforming existing baselines across various base models and benchmarks. Code is available at https://github.com/Tencent/ContextPilot.

Comments10 pages, 6 figures, 5 tables, accepted to EMNLP 2026 (Main Track)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑