CURA:面向计算机使用智能体的可验证运行时警报
CURA: Certified Runtime Alarms for Computer-Use Agents
浏览论文内容
中文总结 AI 辅助
针对计算机使用智能体自报告失效问题,提出仅读取测试环境遥测的外部监控器CURA,可在运行时检测失败,提升任务解决率。
中文摘要 AI 辅助
自报告是部署者拥有的成本最低的监督渠道,但在功能强大的计算机使用智能体(Computer-Use Agents, CUAs)上,它恰恰在监督最为关键的地方失效。在361个OSWorld任务上,我们的流水线(包含只读可行性门、规划器和GUI执行器)达到了82.9的平均任务得分,高于72.4的人类基准,但在71次失败中有64次(90%)以成功声明告终,61次未承认存在阻碍,且在约9100次调用中从未使用明确的失败提示。我们引入了CURA(Certified Runtime Alarms for Computer-Use Agents,面向计算机使用智能体的可验证运行时警报),这是一种仅读取测试环境可见遥测数据的外部监控器,无需模型内部信息、额外的大语言模型(LLM)调用或提示词更改,能将运行轨迹转化为具有可验证误报控制的序列测试。在α=0.10时,其CUSUM警报在终止前中位数31步检测到42.3%的失败,实现的误报率为0.066,且风险可在首次行动前部分解决(门探测的AUROC为0.69)。回顾性分析中,该组合达到0.828的AUROC(内部折减下限为0.802),但它与总标记基线的差距不显著(差值为+0.026,p=0.101);这种差距体现在在线场景中,CURA在匹配的可验证预算下召回率更高:在α=0.10时为0.41,而总标记基线为0.34;在α=0.20时为0.56,总标记基线为0.38。经警报门控的执行中监督在70次失败中恢复了23次,同时将前沿监督器用于38次,形成了可部署的级联方案,平均得分达86.8,完全解决率为84.5%(305/361)。该证书仅限制误报。我们还报告了行为监控无信息的场景。
英文摘要
Self-report is the cheapest oversight channel a deployer has, and on capable computer-use agents (CUAs) it fails precisely where oversight matters. On 361 OSWorld tasks our pipeline, a read-only feasibility gate, a planner, and a GUI executor, reaches a mean task score of 82.9, above the 72.4 human reference, yet 64 of its 71 failures (90%) end with a success claim, 61 acknowledging no blocker, and the explicit failure affordance is never used in roughly 9,100 calls. We introduce CURA (Certified Runtime Alarms for Computer-Use Agents), an external monitor that reads only harness-visible telemetry, with no model internals, extra LLM calls, or prompt changes, and turns the running trajectory into a sequential test with certified false-alarm control. At alpha = 0.10 its CUSUM alarm detects 42.3% of failures a median of 31 steps before termination at a realized false-alarm rate of 0.066, and risk is partly resolvable before the first action (gate probe, 0.69 AUROC). Retrospectively the composite reaches 0.828 AUROC (fold-internal floor 0.802), but its margin over a total-token baseline is not significant (Delta = +0.026, p = 0.101); the separation is online, where CURA recalls more at matched certified budgets: 0.41 versus 0.34 at alpha = 0.10, 0.56 versus 0.38 at alpha = 0.20. Alarm-gated mid-execution oversight recovers 23 of 70 failures while spending a frontier overseer on 38, giving a deployable cascade at mean score 86.8 and 84.5% full-solve (305 of 361). The certificate bounds false alarms only. We also report where behavioral monitoring is uninformative.
发表机构
- University of Illinois Chicago(伊利诺伊大学芝加哥分校)
- Intel Labs(英特尔实验室)
- Capital One AI Labs(Capital One AI实验室)
机构由 AI 辅助整理,请以论文原文为准。