Designing for Doubt: The Case for Informed Abstention in Autonomous Agents
基准测试无法衡量的:论自主智能体弃权能力的评估
机构 * Brown University(布朗大学)
专题命中 Agent评测 :autonomous agent(title,abstract);agent(abstract);agentic(abstract);分类 cs.AI
AI总结 本文指出自主智能体基准测试忽视弃权能力,提出合规偏差概念,并引入弃权场景分类和评估协议,实验表明安全-可用性权衡是可调的。
Comments Accepted to AIES 2026, this is an updated version to the RLEval workshop paper