arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EvoIn:弥合进化与内化以实现智能体微调

EvoIn: Bridging Evolution and Internalization for Agent Fine-Tuning

Shihan Dou, Shaofan Liu, Zhonghang Lu, Jiahang Lin, Shichun Liu, Binghai Wang, Jiajie Jin, Guanting Dong, Tao Gui, Qi Zhang, Xuanjing Huang

arXiv 2609.35290首次发表:更新:

发表机构

Fudan University; Renmin University of China(复旦大学; 中国人民大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EvoIn通过进化并内化决策程序,提升智能体微调效果,在领域内外分别提高通过率10.9和9.2个百分点。

AI 中文摘要

近期工作通过联合进化智能体的外部工具和模型来探索改进智能体的方法,但往往采取一种“大杂烩”式的方法,将新工具、新决策程序以及模型对进化后外部工具的适应捆绑在单一的智能体改进概念之下。在本文中,我们转而研究智能体如何改进其决策程序。具体而言,我们提出EvoIn,一个弥合进化与内化的智能体微调框架。EvoIn首先分析智能体的执行轨迹,通过临时在外部工具中实例化新的决策程序来进化并验证这些程序。经过验证的程序引导智能体生成改进的推理轨迹。随后,这些轨迹被重写为自包含的推理轨迹,去除对外部工具指令的显式引用,同时将所诱导的决策逻辑表达为模型自身的推理。最后,EvoIn在重写后的轨迹上微调模型,内化这些程序,使得改进后的决策在推理时无需依赖进化后的外部工具也能持续存在。我们在多个基准上评估EvoIn,发现它能够持续使智能体学习到更强的决策程序,在领域内通过率提高10.9个百分点,在领域外提高9.2个百分点。结果进一步表明,内化的决策程序能够泛化到未见过的任务。案例研究表明,智能体可以学会在解决任务之前决定如何解决,例如通过检查文档长度来选择是完整阅读还是搜索。EvoIn还具有广泛的适用性,在另一个模型家族上展现出持续的改进。

英文摘要

Recent work has explored improving agents by jointly evolving their harnesses and models, but often takes a ''potpourri'' approach that bundles together new tools, new decision-making procedures, and model adaptation to the evolved harness under a single notion of agent improvement. In this paper, we instead investigate how agents can improve their decision-making procedures. In particular, we propose EvoIn, an agent fine-tuning framework that bridges evolution and internalization. EvoIn first analyzes agent execution traces to evolve and validate new decision-making procedures by temporarily instantiating them in the harness. The validated procedures guide the agent to generate improved reasoning traces. These traces are then rewritten into self-contained reasoning traces, removing explicit references to harness instructions while expressing the induced decision logic as the model's own reasoning. Finally, EvoIn fine-tunes the model on the rewritten traces, internalizing these procedures so that the improved decision-making persists without the evolved harness at inference time. We evaluate EvoIn on diverse benchmarks and find that it consistently enables agents to learn stronger decision-making procedures, raising the pass rate by 10.9 points in-domain and by 9.2 points out-of-domain. Results further show that the internalized decision procedures generalize to unseen tasks. Case studies show that agents can learn to decide how to solve a task before solving it, for example by checking a document's length to choose between reading it in full and searching it. EvoIn is also broadly applicable, showing consistent improvements on another model family.

Comments36 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑