AI 中文总结
研究计算机使用模型对桌面GUI转换的理解,引入桌面增量基准测试(DDB)及其实例、任务,评估多个模型家族,发现排序不饱和,任务上下文对匹配有影响,单动作推断家族更难,DDB填补诊断层,助力改进桌面CUA。
AI 中文摘要
计算机使用代理(CUA)越来越多地通过桌面GUI来完成长期任务。当前基准主要衡量最终任务的成功或单帧定位。两者都无法确定模型是否能够重建由动作产生的因果、与任务相关的转换,而这对于拒绝过时观察、验证进度和从失败中恢复至关重要。这很困难,因为推理、远程输入、应用渲染和屏幕截图捕获是异步的。我们引入了桌面增量基准测试(DDB),这是一个离线步骤级基准测试,有来自约15个应用程序和50个任务域的新颖多应用Linux轨迹的2013个人工验证实例。DDB轨迹通过两个互补任务针对三个失败维度——状态验证、源跟踪和上下文感知控制。我们评估了8个封闭和开源模型家族在32个排序和16个单动作设置下的情况,观察到了一致的差距。排序仍未饱和,最佳非诱饵和诱饵精确匹配率分别为65.1%和65.7%。任务上下文将诱饵识别提高了6.9个百分点,但将非诱饵精确匹配降低了2.2个百分点;错误分析揭示了对呈现的A - B - C顺序的系统复制。单动作结果表明,推断动作家族比定位它更难:点击F1为0.96,而拖动为0.76,同时识别出的拖动通常定位良好。因此,DDB通过填补GUI定位和最终任务成功之间缺失的诊断层,补充了端到端基准测试,能够有针对性地改进桌面CUA的验证、可靠性和恢复。
英文摘要
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates whether a model can reconstruct the causal, task-relevant transition produced by an action- crucial for rejecting stale observations, verifying progress, and recovering from failure. This is difficult because inference, remote input, app rendering, and screenshot capture are asynchronous: the next observation may be delayed, occluded, transient, or unrelated, then misread as progress and carried into subsequent planning. We introduce Desktop-Delta Bench (DDB), an offline step-level benchmark with 2,013 human-verified instances from novel, multi-app Linux trajectories across ~15 applications and 50 task domains. DDB trajectories targets 3 failure dimensions- state verification, source tracking, and context-aware control- through 2 complementary tasks: 463 3-frame temporal-ordering instances, including 105 with a cross-trajectory decoy, and 1,550 before-after pairs labeled from 5 actions + its payload. We evaluate 8 closed and open-source model families across 32 ordering and 16 single-action settings, observing consistent gaps. Ordering remains unsaturated: best non-decoy and decoy exact-match rates are 65.1% and 65.7%. Task context improves decoy identification by 6.9 percentage points but reduces non-decoy exact match by 2.2 points; error analysis reveals systematic copying of the presented A-B-C order. Single-action results show that inferring the action family is harder than locating it: click F1 is 0.96 vs, 0.76 for drag, while recognized drags are generally localized well. DDB, thus, complements end-to-end benchmarks by filling the missing diagnostic layer between GUI grounding and final task success, enabling targeted improvements to desktop CUA verification, reliability, and recovery.