arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18066cs.AIcs.CLcs.LG

自改进智能体的脆弱性:方差、任务顺序与规格不足

On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification

  • Salesforce AI Research(Salesforce AI研究院)

机构由 AI 辅助整理,请以论文原文为准。

Qinyuan Ye, Yu Li, Yada Pruksachatkun, Jiaxin Zhang, Chien-Sheng Wu

中文总结 AI 辅助

本研究通过多轮运行与打乱任务顺序的实验,揭示当前基于记忆的自改进智能体存在脆弱性,发现其性能受方差与任务顺序影响,还指出规格不足是相关诱因,呼吁采用更严格评估方案与人工监督机制。

中文摘要 AI 辅助

基于记忆的自改进智能体(即从在线任务流中学习,通过维护文本记忆库随时间提升性能的智能体)在近年研究中展现出巨大潜力,但这类方法的可靠性方面却被严重忽视。本研究对两种基于记忆的方法展开全面重新评估,沿两个维度拓展评估范围:一是纳入多次运行以量化方差,二是随机打乱任务顺序以探究任务顺序的影响。通过这些实验,我们得出两项暴露当前方法脆弱性的观察结果:其一,在复杂环境与多步任务中,智能体评估本就存在固有噪声,在此基础上叠加自改进循环会进一步放大这种噪声;其二,智能体的提升高度依赖任务顺序,过往研究常采用默认任务顺序,该顺序会施加隐性课程,成为成功的隐藏前提。为更好理解这种脆弱性,我们手动检查智能体的记忆并假设任务与环境规格不足是导致脆弱性的原因,我们通过在记忆构建过程中纳入能实现更优规格的信息(如详细评分标准与环境反馈)来验证该假设。尽管这些新增信息在一定程度上缓解了过往实验中的性能下降,但仍存在显著差距,表明还有其他未明确的因素导致这种脆弱性。展望未来,本研究主张采用更严格的自改进智能体评估方案,即报告多次运行的结果并在挑战性条件下对其进行压力测试;此外,我们关于规格不足的发现呼吁构建能实现有效人工监督的系统与接口,防止智能体以不可预见的方式失效。

英文摘要

Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods have been critically overlooked. In this work, we conduct a comprehensive re-evaluation of two memory-based methods, broadening the scope of evaluation along two axes: (1) including multiple self-improving runs to quantify variance, and (2) shuffling the tasks to investigate the effect of task order. Through these experiments, we make two observations that expose the fragility of current methods: First, agent evaluation is inherently noisy in complex environments and on multi-step tasks, and stacking a self-improving loop on top can further amplify this noise. Second, the agent's improvement is highly dependent on task order. Prior works often adopt default orderings that impose an implicit curriculum, acting as a hidden prerequisite for success. To better understand this fragility, we manually examine the agents' memory and hypothesize that task and environment underspecification contribute to this fragility. We validate this hypothesis by incorporating information that enables better specification, such as detailed rubrics and environment feedback, into the memory construction process. While this added information partially closes the performance degradation in previous experiments, significant gaps still remain, suggesting that other uncharacterized factors contribute to this fragility. Looking ahead, our work advocates for more rigorous evaluation protocols for self-improving agents by reporting results across multiple runs and stress-testing them under challenging conditions. Moreover, our findings on underspecification call for systems and interfaces that enable effective human oversight, preventing agents from failing in unforeseeable ways.

补充信息

↑