arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MemoryWalker:停止在智能体从未见过的上下文上进行训练

MemoryWalker: Stop Training Agents on Contexts They Never Saw

Zinco J, Xunjie Zhu, Shen Huang, Zhenyi Wang, Pengjun Xie, Jieping Ye

arXiv 2609.00865首次发表:更新:

发表机构

Token Foundry, Alibaba Group(Token Foundry,阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体压缩上下文训练引发的树状学习对象问题,提出LogitTree、4D注意力掩码及SDCC方法,在七个网络搜索基准上验证SDCC可缩小训练-rollout差距并提升奖励。

AI 中文摘要

Claude Code、Qwen-Agent等生产级智能体工具在rollout过程中会压缩上下文,但在压缩条件下训练会引发条件问题:每一次上下文驱逐都会使有效历史产生分支,因此学习对象是一棵树而非序列。现有的线性化方法要么保留最右侧路径,引发时间旅行泄漏,要么重播深度优先遍历,引发训练-推理不匹配。我们引入两种精确的、梯度等价的修正方法:LogitTree,一种分段K向前遍历;以及一种打包的4D注意力掩码。LogitTree需要K+1次反向传播;4D掩码需要自定义内核和白盒驱逐记录。我们还提出SDCC(Self-Distillation for Conditioning Consistency,条件一致性自蒸馏),一种单反向传播的变分松弛方法。在每次驱逐时,它最小化压缩学生模型与停止梯度教师在重构的驱逐前前缀上的前向KL散度。每个节点的残差KL为ε_KL,给出了训练-部署总变差差距的O(√ε_KL)边界。SDCC也适用于黑盒工具。在包含TC-RAG、AgentFold、MemexRL、Claude Code和OpenCode的七个网络搜索基准上,朴素训练会扩大训练-rollout对数概率差距,尤其是在驱逐密集的批次中。精确方法保持无压缩的基准水平,而SDCC大幅缩小了该差距,具有更低的logit漂移和更高的rollout奖励。

英文摘要

Production agent harnesses such as Claude Code and Qwen-Agent compress context during rollout, but training under compression creates a conditioning problem: every eviction branches the effective history, so the learning object is a tree rather than a sequence. Existing linearizations either retain the rightmost path, causing time-travel leakage, or replay a depth-first traversal, causing train-inference mismatch. We introduce two exact, gradient-equivalent corrections: LogitTree, a segmented K-forward traversal, and a packed 4D attention mask. LogitTree requires K+1 backward passes; the 4D mask requires a custom kernel and white-box eviction records. We also propose SDCC (Self-Distillation for Conditioning Consistency), a single-backward-pass variational relaxation. At each eviction, it minimizes forward KL between the compressed student and a stop-gradient teacher on the reconstructed pre-eviction prefix. A residual per-junction KL of epsilon_KL gives an O(sqrt(epsilon_KL)) bound on the train-deployment total-variation gap. SDCC also applies to black-box harnesses. On seven web-search benchmarks with TC-RAG, AgentFold, MemexRL, Claude Code, and OpenCode, naive training inflates the train-rollout log-probability gap, especially on eviction-heavy batches. The exact methods stay at the no-compression floor, and SDCC substantially closes the gap, with lower logit drift and higher rollout rewards.

CommentsYour Memory-Compressing Harness Makes Training and Inference Inconsistent

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑