AI 中文总结
针对长视界智能体的上下文增长问题,提出MemOPD方法通过记忆状态对齐解决同策略蒸馏中的状态错位,在提升F1的同时加快训练计算速度,代码公开。
AI 中文摘要
长视界智能体在交互过程中会积累不断增长的上下文,这会损害其性能和稳定性。紧凑记忆通过压缩和重写模型调用之间保留的历史来缓解该问题。学习保留什么内容通常依赖于近端策略优化(PPO)及最终任务奖励,但稀疏奖励无法为单个记忆更新提供充足指导。这一局限性催生了同策略蒸馏(OPD),它为学生的 rollout 提供密集的教师监督。为使此类监督有效,教师必须在动作生成时的同一状态下评估每个采样动作。然而,记忆压缩过程中执行的上下文重写会破坏这种对齐:当采样的响应被保留并重新编码以供后续调用时,将交互展平为持久历史可能导致教师在学生 rollout 期间从未访问过的状态下对动作进行评分。因此,该动作按来源属于同策略,但不一定按状态属于同策略。为此,我们提出记忆对齐同策略蒸馏(MemOPD):MemOPD 记录每个模型调用的输入和采样输出,恢复其原始 token 位置和因果可见性,并打包重构的调用以实现高效的教师评分。教师在采样动作位置提供全词汇监督,同时 PPO 保留最终任务目标。实验验证了多次上下文更新下的状态对齐,显示在匹配对照中,其 F1 比持久历史教师评分提高了7.0%;总体而言,MemOPD-3B 的 F1 比 PPO 最高提升416.2%,且打包操作在训练期间的 actor 计算中实现了最高1.63倍的加速。该工作的代码可公开获取:this https URL
英文摘要
Long-horizon agents accumulate growing contexts during interaction, impairing performance and stability. Compact memory mitigates this problem by compressing and rewriting the history retained between model invocations. Learning what to retain typically relies on proximal policy optimization (PPO) with final task rewards, but sparse rewards provide little guidance for individual memory updates. This limitation motivates on-policy distillation (OPD), which supplies dense teacher supervision on student rollouts. For such supervision to be valid, the teacher must evaluate each sampled action under the same state in which it was generated. However, the context rewriting performed during memory compression can break this alignment. When sampled responses are retained and re-encoded for later invocations, flattening the interaction into a persistent history may cause the teacher to score the action under a state that the student never visited during rollout. The action therefore remains on-policy by provenance, but not necessarily by state. We therefore propose Memory-Aligned On-Policy Distillation (MemOPD). MemOPD records the inputs and sampled outputs of each model invocation, restores its original token positions and causal visibility, and packs the reconstructed invocations for efficient teacher scoring. The teacher provides full-vocabulary supervision at the sampled action positions, while PPO preserves the final task objective. Experiments verify state alignment across several context updates and show that it improves F1 by 7.0% over persistent-history teacher scoring in a matched control. Overall, MemOPD-3B improves F1 over PPO by up to 416.2%, while packing yields up to a 1.63x speedup in actor computation during training. The code for this work is publicly available at: https://github.com/TPssp/MemOPD.