HyMem:基于信息隔离的长程智能体分层上下文管理
HyMem: Hierarchical Context Management for Long-Horizon Agents via Information Isolation
浏览论文内容
中文总结 AI 辅助
针对LLM智能体长程任务上下文杂乱导致推理性能下降的问题,提出分层上下文管理框架HyMem,通过功能分层隔离推理与规划,在GAIA、Browsecomp-plus数据集上取得优于基线的性能。
中文摘要 AI 辅助
大型语言模型(LLM)智能体在复杂长程任务中表现不佳,原因是其上下文会随时间推移变得越来越杂乱。随着交互积累,详细的执行轨迹和中间输出会占据上下文,导致模型难以保留和使用高层规划信息。现有大多数方法通过对单一扁平上下文应用压缩或检索来解决该问题,但这类方法未明确区分不同类型的上下文信息,往往会导致推理性能下降。为应对这一挑战,我们提出HyMem,这是一个将智能体上下文明确划分为不同功能层的分层框架。HyMem按功能组织上下文,将高层规划与执行及复杂分析相分离;其隔离推理模块处理复杂子任务,且不会将中间推理轨迹添加到持久规划上下文中,而其内存管理模块则通过结构化摘要在上下文刷新期间保留任务进度。这些组件减少了冗余上下文积累,保留了任务关键信息,并在有限的上下文窗口内支持连贯的长程推理。在GAIA和Browsecomp-plus数据集上使用DeepSeek-V4进行的实验表明,HyMem的平均Pass@1得分分别为66.7%和61.3%,分别比最强基线高出6.1和4.7个百分点。进一步分析表明,HyMem能有效控制推理上下文的增长,使模型在复杂长程任务中保持专注度和准确性。
英文摘要
Large language model (LLM) agents often perform poorly on complex, long-horizon tasks because their context becomes increasingly cluttered over time. As interactions accumulate, detailed execution traces and intermediate outputs dominate the context, making it difficult for the model to retain and use high-level planning information. Most existing methods address this issue through compression or retrieval applied to a single, flat context, which does not clearly separate different types of context information and often leads to degraded reasoning. To address this challenge, we propose HyMem, a hierarchical framework that explicitly separates the agent's context into distinct functional layers. HyMem organizes context by function to separate high-level planning from execution and complex analysis. Its isolated reasoning module handles complex subtasks without adding intermediate reasoning traces to the persistent planning context, while its memory management module preserves task progress across context refreshes through structured summaries. These components reduce redundant context accumulation, retain task-critical information, and support coherent long-horizon reasoning within a limited context window. Experiments on GAIA and Browsecomp-plus show that, with DeepSeek-V4, HyMem achieves average Pass@1 scores of 66.7% and 61.3%, outperforming the strongest baseline by 6.1 and 4.7 percentage points, respectively. Further analysis indicates that HyMem effectively controls the growth of the reasoning context, allowing the model to maintain focus and accuracy across complex, long-horizon tasks.
发表机构
- Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
- Nanjing Artificial Intelligence Research of IA(IA南京人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。