AI 中文总结
LoongReflect是将反思表述为记忆控制策略的训练框架,通过双信号通道结合全局视角蒸馏与结果导向GRPO优化,在多跳检索增强生成和数学推理任务上提升了搜索智能体的长程反思能力。
AI 中文摘要
大型语言模型智能体越来越依赖长程推理来解决涉及规划、工具使用和记忆的复杂任务。在此类场景中,一项关键能力是反思:评估轨迹进度、识别缺失证据和不可靠的中间状态,并决定是继续、修正还是放弃当前分支。然而,学习有效的反思颇具挑战性,因为反思是在当前分支内局部执行的,而其效用只能通过对最终轨迹结果的贡献来确定。这种局部-全局不匹配使得基于结果的强化学习只能为反思决策提供局部、稀疏且延迟的监督。为解决这些问题,我们提出LoongReflect,这是一个将反思表述为记忆控制策略的训练框架。智能体在可逆轨迹树上运行,使用显式的反思(reflect)和回溯(backtrack)动作。反思会将已验证事实、缺失证据和分支特定风险整合到工作记忆中,而回溯则会从活跃上下文中移除不可靠分支并保留简洁的修正经验。为学习该策略,LoongReflect通过前瞻额外梯度式协调机制结合两种互补信号:快速通道从特权教师中蒸馏全局知情的反思行为,监督范围仅限于反思和回溯标记;慢速通道使用基于结果的GRPO优化完整轨迹,使局部控制决策与最终任务成功对齐。在多跳检索增强生成和数学推理基准上的实验表明,与仅基于结果的强化学习及自蒸馏基线相比,LoongReflect取得了一致的性能提升。
英文摘要
Large language model agents increasingly rely on long-horizon reasoning to solve complex tasks involving planning, tool use, and memory. A critical capability in such settings is reflection: assessing trajectory progress, identifying missing evidence and unreliable intermediate states, and deciding whether to continue, revise, or abandon the current branch. Learning effective reflection, however, is challenging because reflection is performed locally within the current branch, whereas its utility can only be determined by its contribution to the final trajectory outcome. This local-global mismatch makes outcome-based reinforcement learning provide only local, sparse and delayed supervision for reflective decisions. To solve these, we propose LoongReflect, a training framework that formulates reflection as a memory-control policy. The agent operates over a reversible trajectory tree using explicit reflect and backtrack actions. Reflection consolidates verified facts, missing evidence, and branch-specific risks into working memory, while backtracking removes an unreliable branch from the active context and preserves a concise corrective lesson. To learn this policy, LoongReflect combines two complementary signals through a look-ahead, extragradient-style coordination mechanism. A fast channel distills globally informed reflective behavior from a privileged teacher, with supervision restricted to reflection and backtracking tokens. A slow channel optimizes complete trajectories using outcome-based GRPO, aligning local control decisions with final task success. Experiments on multi-hop retrieval-augmented generation and mathematical reasoning benchmarks demonstrate consistent improvements over outcome-only reinforcement learning and self-distillation baselines.
Comments15 pages, 8 figures