arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

APEx:用于自适应深度研究问答的智能体程序经验蒸馏

APEx: Distillation of Agent Procedural Experience for Adaptive Deep Research Question Answering

Jie Ding, Rui Sun, Xinyuan Zhang, Zeyu Zhang, Xin Liu

arXiv 2609.02253首次发表:更新:

发表机构

University of Science and Technology of China; ByteDance Inc.; Chinese Academy of Sciences(中国科学技术大学; 字节跳动公司; 中国科学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

APEx是分层经验利用框架,通过三阶段交替GRPO训练优化Executor、Distiller、Planner模块,经7个基准测试性能超GPT-5.4和最强记忆增强基线,实现自适应深度研究问答。

AI 中文摘要

深度研究智能体借助外部工具增强大型语言模型,通过多轮推理回答复杂、长周期问题。从过往经验学习对持续改进至关重要,但现有方法要么检索冗长的特定任务轨迹,加重决策负担;要么蒸馏与下游策略适应脱节的程序技能。我们提出APEx,这是一种分层经验利用框架,将交互历史组织为实例级轨迹记忆和类别级程序技能,并通过执行器(Executor)、蒸馏器(Distiller)和规划器(Planner)的闭环架构将二者耦合。三个模块通过三阶段交替的GRPO训练范式优化,实现奖励引导的技能蒸馏而非固定提示生成。测试时,蒸馏出的技能作为程序先验,通过技能引导的测试时强化学习用于在线规划器适应,结合无真实值的自我改进与技能对齐正则化,防止策略漂移。在7个基准上的实验表明,APEx实现了SOTA性能,超过GPT-5.4 14.7个点,比最强的记忆增强基线高3.0个点。

英文摘要

Deep research agents augment large language models with external tools to answer complex, long-horizon questions through multi-turn reasoning. Learning from prior experience is crucial for continual improvement, yet existing methods either retrieve verbose task-specific traces that burden decision-making, or distill procedural skills that remain decoupled from downstream policy adaptation. We propose APEx, a hierarchical experience utilization framework that organizes interaction history into instance-level trajectory memories and category-level procedural skills, and couples them through a closed-loop architecture of Executor, Distiller, and Planner. The three modules are optimized via a three-stage alternating GRPO training paradigm, enabling reward-guided skill distillation rather than fixed-prompt generation. At test time, distilled skills serve as procedural priors for online Planner adaptation through skill-guided test-time reinforcement learning, allowing ground-truth-free self-improvement with skill-alignment regularization to prevent policy drift. Experiments on 7 benchmarks demonstrate that APEx achieves state-of-the-art performance, surpassing GPT-5.4 by 14.7 points and the strongest memory-augmented baseline by 3.0 points.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑