发表机构
University of Dhaka; University of Maryland, Baltimore County(达卡大学; 马里兰大学巴尔的摩县分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出BACKDROP基准,通过四种日常危害测试智能体在动态世界中的鲁棒性,发现平均通过率从69.5%降至31.3%,揭示现实性能上限。
AI 中文摘要
智能体基准测试在静止不变的世界中测试智能体。而实际部署的智能体工作在其他人类也会改变的世界中。有人给智能体发短信要求将钱转往别处,或者订单确认信息要求它回复门禁密码。我们提出BACKDROP,用于探究智能体在干净世界中的能力有多少能在这种世界中存续。BACKDROP选取一项任务及智能体的执行环境,并在其世界中设置四种日常危害,分别单独出现和同时出现。指令和正确的最终状态保持不变。每种危害提出一个问题。权威性:来自他人的消息是否会覆盖用户的指令?注入:记录中植入的文本是否会误导智能体?边界:请求是否会将其拉入未被授予权限的应用?故障:写入失败且未说明是否成功时,智能体是否会在重试前进行检查?在3678个变体和16个模型中,当四种危害同时存在时,平均通过率从69.5%降至31.3%;最强模型下降幅度最大(Claude Fable 5.1从96.6%降至56.0%)。智能体已学会抵抗注入文本,但常常遵循他人的其他未经授权请求。在四种危害同时存在且仅统计植入文本到达智能体的运行中,智能体在46.4%的运行中遵循了他人的消息,在20.3%的运行中遵循了注入文本。这一差距在所有16个模型中一致存在。BACKDROP形式化了这些差距,并展示了智能体在任务世界中的得分是现实世界表现的上限。
英文摘要
Agent benchmarks test agents in worlds that stay still. Deployed agents work in worlds that other people also change. Someone texts the agent to send the money elsewhere or an order confirmation asks it to reply with a door code. We present BACKDROP, which asks how much of an agent's capability in a clean world survives in such a world. BACKDROP takes a task along with the agents execution environment, and plants four everyday hazards in its world, one at a time and all together. The instruction and the correct end state stay the same. Each hazard asks one question. Authority: does a message from another person override the user? Injection: does text planted in a record redirect the agent? Boundary: does a request pull it into an app it was not given? Fault: after a write fails without saying whether it landed, does the agent check before it retries? Across 3,678 variants and 16 models, , the average pass rate falls from 69.5% to 31.3% once all four hazards are present; the strongest models fall furthest (Claude Fable 5.1 from 96.6% to 56.0%). Agents have learned to resist injected text but often follow other unauthorized requests of other people. With all four hazards present, and counting only runs where the planted text reached the agent, agents followed another person's message in 46.4% of runs and injected text in 20.3%. The gap is consistent throughout all 16 models. BACKDROP formalizes these gaps and shows how an agent's score in a task's world is a ceiling on real-world performance.
CommentsSubmitted to ICLR 2027