arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33870cs.AI

当成功策略失效:终端智能体对环境新颖性的适应

When Successful Strategies Fail: Adaptation to Environmental Novelty in Terminal Agents

Janvijay Singh, Vaishnavi Shrivastava, Dilek Hakkani-Tur, Ece Kamar, Asli Celikyilmaz

首次发表
浏览论文内容

中文总结 AI 辅助

本文研究LLM智能体在环境假设失效时的适应能力,提出AGNI流水线生成环境新颖性,发现智能体存在适应差距,并证明后训练可提升适应性和基础性能。

中文摘要 AI 辅助

LLM智能体通过自主与环境交互,日益能够解决长期任务。在此过程中,它们的策略依赖于对环境的假设:存在哪些资源和工具、它们位于何处以及它们如何表现。当这些假设不再成立时,可靠的智能体必须在追求相同目标的同时检测变化并适应。我们通过环境新颖性来研究这种适应能力:一种保持任务目标不变,同时使原本成功轨迹所依赖的假设失效的变化。我们引入了AGNI,一个自动化流水线,用于提取与轨迹相关的假设、注入有针对性的环境变化,并验证由此产生的新颖任务仍然可解。在三个终端基准测试中,AGNI生成了涵盖资源、接口、约束和执行语义的多样化新颖性。评估多个LLM智能体揭示了基础任务和新颖任务之间存在显著的适应差距。轨迹分析表明,智能体经常遇到变化的证据,但未能诊断其原因并修订其策略。最后,针对环境新颖性的后训练提高了对保留的新颖任务的适应能力,同时也提高了基础任务的性能。我们的结果突显了任务能力与适应能力之间的差距,并促使环境变化作为智能体训练和评估的核心维度。

英文摘要

LLM agents increasingly solve long-horizon tasks by autonomously interacting with their environment. In doing so, their strategies rely on assumptions about that environment: which resources and tools exist, where they are located, and how they behave. When these assumptions no longer hold, reliable agents must detect the change and adapt while pursuing the same goal. We study this adaptation capability through environmental novelty: a change that keeps the task objective fixed while invalidating an assumption underlying an otherwise successful trajectory. We introduce AGNI, an automated pipeline that extracts trajectory-relevant assumptions, injects targeted environmental changes, and validates that the resulting novel tasks remain solvable. Across three terminal benchmarks, AGNI produces diverse novelties spanning resources, interfaces, constraints, and execution semantics. Evaluating multiple LLM agents reveals a substantial adaptation gap between base and novel tasks. Trajectory analysis suggests that agents often encounter evidence of the change but fail to diagnose its cause and revise their strategy. Finally, post-training for environmental novelty improves adaptation to held-out novel tasks while also improving performance on base tasks. Our results highlight a gap between task competence and adaptive capability and motivate environmental variation as a core dimension of agent training and evaluation.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
  • Microsoft Research AI Frontiers(微软研究院AI前沿)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑