arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

发现AppWorld和WorkArena任务验证器中的盲点

Finding Blind Spots in AppWorld and WorkArena Task Verifiers

Richard Abrich

arXiv 2610.09142首次发表:更新:

发表机构

OpenAdapt.AI (MLDSAI Inc.)(OpenAdapt.AI(MLDSAI公司))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

通过源代码知情的变异测试审计AppWorld和WorkArena验证器,发现其存在接受错误效果和拒绝有效执行的盲点,并报告了具体案例与拒绝普查结果。

AI 中文摘要

基于执行的任务验证器决定智能体是否成功。我们使用源代码知情的变异测试审计了已发布的AppWorld和WorkArena验证器。主要审计从不修改已发布的检查器。在AppWorld中,重复非幂等写入会创建额外记录,同时保留所有被检查的字段值。验证器接受了来自五个合格生成器中的两个的所有三个任务变体:6/15构造效果。在普查后应用于检查器副本的基数补丁使所有六个单元失败,同时保留有效对照。在WorkArena中,我们前瞻性地重新运行了23个额外字段候选,这些候选先前被选为检查器通过结果。独立的Table API回读确认了21个非默认持久化值,而所有23个都获得通过。两个请求的字符串是存储默认值的别名。21个确认的错误效果跨越三个表单模板。这些选定的案例在审计协议下确认了错误效果;它们不估计总体率。没有其他构造产生独立确认的错误接受。其他检查器通过的案例是效果正确的退化。我们单独报告零通过家族,因为保留的证据不同。在固定的意图交换网格中,检查器在2,689个非对角线执行上返回无通过。这是一个拒绝普查:57个WorkArea单元使用会话范围的证据;其他2,632个缺乏分类的拒绝原因和独立的目标真实值。每个增量在其自身单元被评分之前指定。补充材料随OpenReview提交附带构造语法、证据、内容分阶段谱系和计数复现器。

英文摘要

Execution-based task verifiers decide whether an agent succeeded. We audit shipped AppWorld and WorkArena verifiers with source-informed mutation tests. The main audit never modifies a shipped checker. In AppWorld, duplicating a non-idempotent write creates an extra record while preserving every checked field value. The verifier accepts all three task variants from two of five eligible generators: 6/15 constructed effects. A cardinality patch applied to checker copies after the census makes all six cells fail while preserving valid controls. In WorkArena, we prospectively rerun 23 extra-field candidates selected for earlier checker-PASS outcomes. Independent Table API readback confirms nondefault persisted values in 21, while all 23 receive PASS. Two requested strings are aliases of stored defaults. The 21 confirmed wrong effects span three form templates. These selected cases confirm wrong effects under the audit's protocol; they do not estimate a population rate. No other construction produces an independently confirmed false accept. Other checker-PASS cases are effect-correct degeneracies. We report zero-PASS families separately because retained evidence differs. In fixed intent-swap grids, the checkers return no PASS on 2,689 off-diagonal executions. This is a rejection census: 57 WorkArena cells use session-scoped evidence; the other 2,632 lack classified rejection causes and independent target ground truth. Each increment is specified before its own cells are scored. A supplement accompanies the OpenReview submission with the construction grammar, evidence, content-bound stage lineage and count reproducer.

Comments14 pages. Accepted as a poster at the NeurIPS 2026 workshop "Who Verifies the Agents? Toward Reliable Agent Development". OpenReview paper and supplement: https://openreview.net/forum?id=uR6EWsGE48

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑