arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.11677cs.SEcs.AI

Ecdysis:面向LLM智能体的运行时框架高效有效训练

Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents

  • Chengdu Institute of Computer Applications, Chinese Academy of Sciences(中国科学院成都计算机应用研究所)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Beijing Institute of Technology(北京理工大学)
  • Beijing University of Technology(北京工业大学)
  • Yangtze Delta Region Institute of Tsinghua University, Zhejiang(浙江清华大学长三角研究院)
  • Jiaxing Key Laboratory of Artificial Intelligence and Cyber Resilience(嘉兴人工智能与网络安全重点实验室)

机构由 AI 辅助整理,请以论文原文为准。

Ruiqing Yue, Yu Cui, Zhuoyu Sun, Sicheng Pan, Xianhong Xue, Tingyu Li, Ting Li, Wenzhuo Zhu, Yi Chen, Yifei Liu, Baohan Huang, Zhe Cui, Haibin Zhang, Cong Zuo

AI总结:

针对LLM智能体运行时框架演化中失败诊断缺失与训练开销大的问题,提出Ecdysis框架,通过跨实例失败聚合与协作精炼区分模型缺陷与框架缺陷,实现1.84倍训练加速和18.56%准确率提升。

AI中文摘要:

自我演化的运行时框架可以显著提升大语言模型(LLM)智能体的能力,并为优化智能体执行提供了一种有前景的范式。现有的框架演化方法通常依赖迭代搜索,基于任务实例的执行反馈,反复评估和修改候选框架。虽然这种范式能够实现框架的持续优化,但由于重复的智能体执行和代码修改,它带来了大量的时间开销,并且可能过拟合到已观测的任务和特定的失败模式,导致对未见任务的泛化能力下降。我们识别出缺乏原则性的失败诊断是框架演化的关键瓶颈:一个观测到的失败可能反映模型特定的缺陷或系统性的框架缺陷,而直接针对单个失败进行优化可能导致不必要的模型特定适配。因此,我们提出了Ecdysis,一个高效且有效的框架,它区分模型特定适配与框架级修复,并通过识别跨任务重复出现的失败模式,将适应偏向于系统性的框架缺陷。Ecdysis采用批量级别的跨实例失败聚合范式,联合分析来自多个任务实例的失败证据,并进一步引入失败驱动的协作精炼来诊断失败原因并迭代精炼框架修改规范。通过将跨实例失败分析与多角色诊断相结合,Ecdysis能够以更低的训练时间实现更有效的框架演化。实验表明,与现有的框架演化方法相比,Ecdysis在框架训练上实现了高达1.84倍的加速,同时将所得框架的推理准确率提高了18.56%。

英文摘要:

Self-evolving runtime harnesses can substantially improve the capabilities of large language model (LLM) agents and provide a promising paradigm for optimizing agent execution. Existing failure-driven approaches often treat observed agent failures as direct evidence for harness modification. A key challenge in failure-driven harness evolution is that observed failures can reflect either limitations of the underlying model or systematic deficiencies of the harness. Directly optimizing against individual failures can therefore induce model-specific accommodation and impair generalization across tasks and models. We study whether failure evidence accumulated across task instances can provide a more reliable signal for harness training. Our key insight is that failures recurring across distinct tasks provide stronger inductive evidence for systematic harness deficiencies than isolated failures. Based on this insight, we propose Ecdysis, which aggregates failure evidence across task instances before promoting recurring failure patterns into persistent harness evolution, biasing evolution toward repairs that are more likely to generalize beyond individual model behaviors. Ecdysis further employs collaborative failure analysis to refine modification specifications, trading additional evolution-time reasoning for improved modification quality. Across multiple LLMs and benchmarks, Ecdysis improves the reasoning accuracy of evolved harnesses by 18.56% over existing harness evolution while achieving up to 1.84x faster harness training. Ecdysis also enables more data-efficient training. Fine-grained analysis shows that Ecdysis reduces model-specific accommodation during evolution, while the resulting harnesses exhibit stronger cross-LLM generalization and lower inference-time token consumption.

↑