arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ASCENT:通过验证经验的自蒸馏实现长程智能体的在线测试时训练

ASCENT: Online Test-Time Training of Long-Horizon Agents via Self-Distillation of Verified Experience

Haodong Lu, Dong Gong

arXiv 2610.05303首次发表:更新:

发表机构

University of New South Wales (UNSW Sydney)(新南威尔士大学(悉尼新南威尔士大学))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ASCENT提出一种在线测试时训练方法,通过自蒸馏验证经验更新LLM权重,提升长程任务成功率与交互效率,并泛化到新场景。

AI 中文摘要

大型语言模型(LLM)智能体通过多次推理-行动轮次解决长程任务,并在终止时获得一个验证信号。部署中的智能体面对相关任务的流式输入,使其轨迹成为改进的自然资源。上下文适应智能体将反思、记忆或技能以文本形式存储,因此重用依赖于检索正确的经验以及冻结策略执行该经验。我们研究在线智能体测试时训练(OaTTT),该训练在部署期间基于智能体自身的执行轨迹训练LLM的权重。智能体对每个任务执行一次,在流上单次遍历,执行轨迹及其验证结果是用于跨任务持久权重更新的唯一学习信号。直接模仿或强化单次尝试生成的标记会破坏策略的稳定性。我们引入ASCENT(Agentic Self-distillation for Cross-task EvolutioN at Test-time),它转而自蒸馏验证经验。LLM的稳定版本,即其冻结的初始副本,接收验证轨迹作为特权信息,并沿该轨迹以事后视角预测下一标记分布。将这些分布蒸馏到持久的LoRA快速权重中,为后续任务更新智能体,无需外部参考解决方案或更强的教师模型。通过进一步移除无效行动轮次,ASCENT蒸馏增强的特权经验以实现更高效的执行。我们刻画其群体目标及稀疏结果选择的限制。在ALFWorld、WebShop和AppWorld上,跨不同模型规模,ASCENT随着经验积累提高任务成功率和交互效率,优于在线适应方法,并泛化到未见场景,表明智能体可以在无需单独训练阶段或记忆检索的情况下将验证经验整合到其权重中。项目页面:此https URL

英文摘要

A large language model (LLM) agent solves long-horizon tasks through many reasoning-action turns, with one verification signal at termination. Deployed agents face streams of related tasks, making their trajectories a natural resource for improvement. In-context adaptation agents store reflections, memories, or skills as text, so reuse depends on retrieving the right experience and on a frozen policy executing it. We study Online Agentic Test-Time Training (OaTTT), which trains the LLM's weights on its own execution trajectories during deployment. The agent executes each task once, in one pass over the stream, and the executed trajectory with its verification result is the only learning signal for weight updates that persist across tasks. Directly imitating or reinforcing the generated tokens of this single attempt destabilizes the policy. We introduce ASCENT (Agentic Self-distillation for Cross-task EvolutioN at Test-time), which instead self-distills verified experience. A stable version of the LLM, its frozen initial copy, receives the verified trajectory as privileged information and predicts next-token distributions along it with this hindsight. Distilling them into persistent LoRA fast weights updates the agent for later tasks, without an external reference solution or stronger teacher. By further removing invalid-action turns, ASCENT distills enhanced privileged experience for more efficient execution. We characterize its population target and the limits of sparse outcome selection. Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes, showing that an agent can consolidate verified experience into its weights without a separate training phase or memory retrieval. Project page: https://artificer-ai-lab.github.io/ASCENT

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑