arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从修订后果中学习:基于事后元经验蒸馏的自我改进智能体

Learning from Revision Consequences: Hindsight Meta-Experience Distillation for Self-Improving Agents

Qianhan Feng, Zhongzhen Huang, Yakun Zhu, Xiaofan Zhang, Qi Dou

arXiv 2610.07979首次发表:更新:

发表机构

The Chinese University of Hong Kong; Shanghai Jiao Tong University(香港中文大学; 上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对智能体元技能学习受分支结果干扰的问题,提出HMED机制,通过从相同状态重放修订前后元技能并蒸馏元经验,在多个基准上显著提升技能发现性能。

AI 中文摘要

随着智能体通过生成和修订技能而持续改进,发现并精炼这些技能的过程本身也成为一个可学习的对象。任务技能直接作用于任务执行,而元技能则控制智能体如何发现和改进未来的技能;因此,元技能的价值通过其引发的后续搜索过程得以体现。现有方法从观察到的原始技能搜索轨迹和分支结果中改进元技能。然而,分支性能混淆了初始发现状态与生成搜索过程的元技能修订的影响,使得难以刻画特定修订实际改变了什么,并促使更新偏向于受益于有利状态而非改进过程的修订。我们提出了HMED(事后元经验蒸馏),一种为自我改进智能体构建元经验的机制。HMED重新审视修订所源自的已完成事件,并从相同的恢复发现状态重新执行现有和修订后的元技能,从而在共享条件下观察与修订相关的变更。每次比较都被蒸馏为一条元经验,这是一种结构化记录,可被未来更新重用,因此即使最终未被保留的修订也能贡献学习信号。在三个交互式智能体基准以及开源和闭源模型上,HMED始终优于强基线,提升了技能发现性能,将元技能学习从分支结果转向改变改进过程的后果。

英文摘要

As agents continuously improve by generating and revising Skills, the process that discovers and refines those Skills becomes a learnable object in its own right. Task-Skills directly act on task execution, whereas Meta-Skills govern how agents discover and improve future Skills; their value therefore emerges through the subsequent search processes they induce. Existing approaches improve Meta-Skills from observed raw Skill-search trajectories and branch outcomes. However, branch performance entangles the effects of the initial discovery state and the Meta-Skill revision that generated the search process, making it difficult to characterize what a particular revision actually changed, and pushing updates toward revisions that benefit from favorable states rather than those that improve the process. We introduce HMED (Hindsight Meta-Experience Distillation), a mechanism for constructing Meta-Experience for self-improving agents. HMED revisits the completed event from which a revision originates and re-executes the incumbent and revised Meta-Skills from the same restored discovery state, so that the changes associated with the revision can be observed under a shared condition. Each comparison is distilled into a Meta-Experience, a structured record that can be reused by future updates, so that even revisions that are not ultimately retained still contribute a learning signal. Across three interactive agent benchmarks and both open-source and closed-source models, HMED consistently improves Skill discovery performance over strong baselines, shifting Meta-Skill learning beyond branch outcomes toward the consequences of changing the improvement process.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑