arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10059cs.AI

AgentAbstain:基于大语言模型的智能体知道何时不行动吗?

AgentAbstain: Do LLM Agents Know When Not to Act?

Xun Liu, Yi Evie Zhang, Vira Kasprova, Parisa Rabbani, Pardis Sadat Zahraei, Tianyu Zhang, Ali Ebrahimpour-Boroojeny, Varun Chandrasekaran

首次发表
浏览论文内容

中文总结 AI 辅助

研究基于大语言模型的智能体何时应弃权的问题,提出AgentAbstain评估框架及AbstainGen全自动管道,通过配对任务基准测试多种前沿LLMs,发现弃权能力与任务解决能力无关,并识别出失败模式,代码和数据集开源。

中文摘要 AI 辅助

基于大语言模型(LLMs)的智能体系统越来越多地用于自主任务,但现有评估大多关注任务成功,而非智能体是否知道何时弃权。这种差距带来实际风险,如在模糊性、冲突约束或工具故障下,智能体可能执行意外且不可逆转的行动。为填补这一差距,我们提出首个用于智能体弃权的系统评估框架:使用工具的LLM智能体识别何时不行动的校准能力。AgentAbstain核心是基于智能体原生分类法构建的配对任务基准,涵盖8种弃权场景。它包含42个可执行沙盒环境中的263个配对任务,通过对指令、工具或环境状态的受控扰动生成。为扩展配对设计并防止数据污染,我们提出AbstainGen,一个全自动管道。在4种智能体框架中的17个前沿LLMs上测试,最佳智能体(Gemini 3.1 Pro)配对准确率仅59.5%。弃权能力很大程度上独立于一般任务解决能力。我们还识别出事后弃权等失败模式。代码和数据集已开源。

英文摘要

Agent systems based on large language models (LLMs) are increasingly deployed for autonomous tasks, yet existing evaluations mostly focus on task success rather than whether agents know when to abstain. This gap poses real risks: under ambiguity, conflicting constraints, or tool failures, agents may execute unintended and irreversible actions. To close this gap, we present the first systematic evaluation framework for agentic abstention: the calibrated ability of tool-using LLM agents to recognize when not to act. At its core, AgentAbstain is a paired-task benchmark built on an agent-native taxonomy of 8 abstention scenarios across pre-execution reasoning and runtime discovery. It contains 263 paired tasks across 42 executable sandbox environments, where each pair consists of a should-act task and a should-abstain variant produced through a controlled perturbation to the instruction, tool, or environment state. To scale this paired design and resist data contamination, we propose AbstainGen, a fully automated pipeline that synthesizes sandbox environments and generates paired tasks end-to-end, validated by deterministic replay and semantic LLM judges; fresh task instances can be regenerated on demand, and three independent annotators rate 94-98% of sampled tasks as well-designed. Across 17 frontier LLMs in 4 agent harnesses, the best agent (Gemini 3.1 Pro) achieves only 59.5% paired accuracy (correct on both the act and abstain sides of each paired task). More importantly, abstention capability is largely independent of general task-solving capability, indicating that scaling task-solving alone will not close this gap. We further identify failure modes such as post-hoc abstention, in which agents execute irreversible actions before recognizing abstention triggers. Our code and dataset are open-sourced at agentabstain.github.io.

发表机构

  • University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑