发表机构
IBM Software Innovation Lab; IBM Research(IBM软件创新实验室; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM智能体在重复任务中表现不一致的问题,提出自进化框架,通过分析轨迹低一致性步骤并生成准则存入记忆,在AppWorld上将五次全成功率提升16个百分点。
AI 中文摘要
大型语言模型(LLM)驱动的智能体在平均表现上可能准确,但在生产环境中却不可靠,这一差异已被观察到,但在很大程度上仍未得到解决。当同一任务被重复执行五次时,使用GPT-4.1的ReAct智能体在AppWorld基准上,五次运行全部成功的概率仅为53%,尽管其单次运行的平均通过率为77%。我们将这24个百分点的差距称为一致性差距,并认为解决这一差距是可信AI智能体部署的先决条件。我们提出了一种自进化智能体框架,通过识别智能体轨迹中不稳定、低一致性的步骤,并将其转化为情景记忆,供智能体在未来的运行中调用,从而缩小这一差距。其核心是一个一致性分析器,用于精确定位轨迹在执行过程中可能发生翻转的位置及原因,以及一个准则生成器,将诊断结果转化为有针对性的准则,提交至记忆并在未来类似任务的智能体执行中注入。在AppWorld上使用ReAct/GPT-4.1,我们的框架将五次运行全部成功的任务比例提升了16个百分点(同任务评估)和13个百分点(相似任务泛化)。
英文摘要
Large language model (LLM)-powered agents can be accurate on average yet unreliable in production, a discrepancy that has been observed but remains largely unaddressed. When given the same task five times, a ReAct agent on the AppWorld benchmark using GPT-4.1 succeeds in all five runs only 53% of the time, even though its per-run pass rate averages 77%. We call this 24-point shortfall the consistency gap, and we argue that addressing it is a precondition for trustworthy AI agent deployment. We present a self-evolving agent framework that reduces this gap by identifying unstable, low-consistency steps in agent trajectories and converting them into episodic memory the agent can draw on in future runs. At its core is a Consistency Analyzer that pinpoints where and why a trajectory is likely to flip across executions, and a Guideline Generator that converts the diagnosis into targeted guidelines, committed to memory and injected into future agent executions on similar tasks. On AppWorld with ReAct/GPT-4.1, our framework raises the fraction of tasks that succeed in all five runs by +16 points on same-task evaluation and +13 points on similar-task generalization.