arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LiteEvo:面向未见任务泛化的自动化、低成本智能体框架演化

LiteEvo: Automated, Cost-Efficient Harness Evolution for Generalization to Unseen Tasks

Euntae Choi, Sumin Song, Sungjoo Yoo

arXiv 2609.33146首次发表:更新:

AI 中文总结

LiteEvo通过无工具元智能体挖掘轨迹组件并构建版本化库,以轻量方式自动演化智能体框架,在多个基准上以更低成本实现与HarnessX相当或更优的泛化性能。

AI 中文摘要

一个LLM智能体由两部分定义:其模型内部的权重以及围绕其组装而成的组件框架。框架目前仍由人工设计,而自动演化框架的HarnessX则从每个基准测试的人工框架出发,报告其在演化任务上的增益,并为每个基准测试预算1亿至1.75亿个元智能体令牌。我们提出LiteEvo,一种轻量级框架演化算法,其无工具元智能体从智能体轨迹中挖掘可复用组件,将其整理成带版本号的库,并据此组合每一轮的框架,从相同的初始框架开始每个基准测试,且从不提及基准测试名称。在五个智能体基准测试的评分任务上,使用冻结的Qwen3.5-9B进行演化,LiteEvo将pass@2提升了10.5至67.7个百分点,并以平均13.0的更低API成本达到了与HarnessX复现相当或更高的pass@2(平均71.0对67.3)。在训练任务上演化的框架在四个基准测试的未见测试任务上保持了增益,LiteEvo还将Claude Code与Sonnet 4.6的组合提升了1.2至71.4个百分点。

英文摘要

An LLM agent is defined by two things: the weights inside its model and the harness of components assembled around it. Harnesses are still handcrafted, and HarnessX, which evolves them automatically, starts each benchmark from a handcrafted harness, reports gains on the tasks it evolved on, and budgets 100 to 175 million meta-agent tokens per benchmark. We propose LiteEvo, a lightweight harness-evolution algorithm whose tool-free meta-agents mine agent trajectories for reusable components, curate them into a versioned library, and compose each round's harness from it, starting every benchmark from the same neutral harness and never naming the benchmark. Evolving on the graded tasks of five agentic benchmarks with a frozen Qwen3.5-9B, LiteEvo lifts pass@2 by 10.5 to 67.7pp and reaches comparable or higher pass@2 than a reproduction of HarnessX (71.0 against 67.3 on average) at 13.0 lower mean API cost. Harnesses evolved on train tasks keep their gains on unseen test tasks of four benchmarks, and LiteEvo also lifts Claude Code with Sonnet 4.6 by 1.2 to 71.4pp.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑