arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为什么GUI智能体正确但延迟?基于决策时间关键路径的解码,用预编译策略树测试

Why Are GUI Agents Correct but Late? Decode on the Decision-Time Critical Path, Tested with Pre-Compiled Policy Trees

Zihan Dong, Rui Qian, Qishi Zhan, Dongshen Peng, Kaixin Li, Yu Li

arXiv 2607.28399首次发表:更新:

AI 中文总结

针对GUI智能体因决策时间关键路径的自回归解码延迟导致瞬态事件失败的问题,提出AAPT策略树方法,提升了决策窗口内的动作成功率,验证了分支路由为关键瓶颈。

AI 中文摘要

计算机使用智能体常因仅在相关窗口关闭后才生成正确动作,而在瞬态GUI事件上失败。我们将主因确定为决策时间关键路径上昂贵的自回归解码。我们提出自适应预期策略树(Adaptive Anticipatory Policy Trees, AAPT),该方法无需修改底层模型即可消除延迟。在屏幕空闲时段,同一冻结的多模态模型会构建带可观测保护、预授权动作及分支特定截止时间的有界条件策略树,树的大小设置为覆盖模型自身的解码延迟。当事件发生时,轻量观察器会将变化门控帧与准备好的分支匹配,并立即执行对应动作,无需生成新文本。在带预注册端点的配对试验及精确McNemar检验中,AAPT在有争议的决策窗口内将成功率从0.50提升至0.79(p=1.8×10⁻³),且未产生任何错误动作。开环及预测-重规划基线均取得零成功率,因其仍在执行期间进行解码。准备阶段的搜索显示,增益出现在基于延迟的树大小规则所预测的位置,消融实验揭示三项关键要求:快速观察器解码、有效树规划及准确分支路由。一项预注册的 oracle 探针拒绝了我们的初始假设,反而指出分支路由是因果瓶颈。我们进一步在126次配对试验中,于独立通用多模态模型上复现了该效应(p=4.9×10⁻¹³)。在外部基准上,AAPT与反应式基线的整体性能相当,尽管二者优势互补。综合来看,这些结果表明,当候选动作可提前枚举时,AAPT表现最佳;而当无法提前枚举时,反应式执行仍更强。

英文摘要

Computer-use agents often fail on transient GUI events because they produce the correct action only after the relevant window has already closed. We identify the main cause as expensive autoregressive decoding on the decision-time critical path. We propose Adaptive Anticipatory Policy Trees (AAPT), which eliminates this delay without modifying the underlying model. During idle screen periods, the same frozen multimodal model constructs a bounded conditional policy tree with observable guards, pre-authorized actions, and branch-specific deadlines. The tree is sized to cover the model's own decoding latency. When an event occurs, a lightweight observer matches change-gated frames to a prepared branch and immediately executes the corresponding action without generating new text. In paired trials with pre-registered endpoints and exact McNemar tests, AAPT improves the success rate from 0.50 to 0.79 within a contested decision window ($p=1.8\times10^{-3}$), while producing no incorrect actions. Both open-loop and predict-and-replan baselines achieve zero success because they still decode during execution. A preparation-time sweep shows that the gain emerges where the latency-based tree-sizing rule predicts, and ablations reveal three key requirements: fast observer decoding, valid tree planning, and accurate branch routing. A pre-registered oracle probe rejects our initial hypothesis and instead points to branch routing as the causal bottleneck. We further reproduce the effect on an independent general-purpose multimodal model over 126 paired trials ($p=4.9\times10^{-13}$). On an external benchmark, AAPT matches the overall performance of a reactive baseline, although the two methods exhibit complementary strengths. Together, these results suggest that AAPT performs best when candidate actions can be enumerated in advance, whereas reactive execution remains stronger when they cannot.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑