arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

弥合一致性差距:学会保持正确轨道的自进化智能体

Closing the Consistency Gap: Self-Evolving Agents That Learn to Stay on Course

Evelyn Duesterwald, Benjamin Elder, Lilian Ngweta, Shashanka Ubaru, Malgorzata Zimon

arXiv 2609.08832首次发表:更新:

发表机构

IBM Software Innovation Lab; IBM Research(IBM软件创新实验室; IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM智能体在重复任务中表现不一致的问题,提出自进化框架,通过分析轨迹低一致性步骤并生成准则存入记忆,在AppWorld上将五次全成功率提升16个百分点。

AI 中文摘要

大型语言模型(LLM)驱动的智能体在平均表现上可能准确,但在生产环境中却不可靠,这一差异已被观察到,但在很大程度上仍未得到解决。当同一任务被重复执行五次时,使用GPT-4.1的ReAct智能体在AppWorld基准上,五次运行全部成功的概率仅为53%,尽管其单次运行的平均通过率为77%。我们将这24个百分点的差距称为一致性差距,并认为解决这一差距是可信AI智能体部署的先决条件。我们提出了一种自进化智能体框架,通过识别智能体轨迹中不稳定、低一致性的步骤,并将其转化为情景记忆,供智能体在未来的运行中调用,从而缩小这一差距。其核心是一个一致性分析器,用于精确定位轨迹在执行过程中可能发生翻转的位置及原因,以及一个准则生成器,将诊断结果转化为有针对性的准则,提交至记忆并在未来类似任务的智能体执行中注入。在AppWorld上使用ReAct/GPT-4.1,我们的框架将五次运行全部成功的任务比例提升了16个百分点(同任务评估)和13个百分点(相似任务泛化)。

英文摘要

Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a discrepancy that has been observed but remains largely unaddressed. When given the same task five times, a ReAct agent on the AppWorld benchmark using GPT-4.1 succeeds in all five runs only 53% of the time, even though its per-run pass rate averages 77%. We call this 24-point shortfall the consistency gap, and we argue that addressing it is a precondition for trustworthy AI agent deployment. We present a self-evolving agent framework that reduces this gap by identifying unstable, low-consistency steps in agent trajectories and converting them into episodic memory the agent can draw on in future runs. At its core is a Consistency Analyzer that pinpoints where and why a trajectory is likely to flip across executions, and a Guideline Generator that converts the diagnosis into targeted guidelines, committed to memory and injected into future agent executions on similar tasks. On AppWorld with ReAct/GPT-4.1, our framework raises the fraction of tasks that succeed in all five runs by +16 points on same-task evaluation and +13 points on similar-task generalization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑