正确的屏幕,错误的过渡:世界模型作为GUI智能体的验证器
Right Screen, Wrong Transition: World Models as Verifiers for GUI Agents
- Nanyang Technological University(南洋理工大学)
- Institute of Science Tokyo(东京科学大学)
- Sony AI(索尼人工智能)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出无解码器的动作条件世界模型LGWM,在185万真实GUI过渡上训练,可直接预测下一屏幕表示,在RSWT-BENCH上验证GUI智能体安全性,性能优于同类模型,为世界模型的验证角色提供了新思路。
AI中文摘要:
点击“登录”后应出现登录屏幕;点击“查看订单”后出现相同屏幕则属于攻击。因此对于GUI智能体而言,安全性是过渡的属性而非屏幕的属性,仅检查屏幕的监控器可通过复用合法屏幕被规避。判断过渡需要知道动作应产生的预期结果。现有GUI世界模型能提供该预期,但它们以文本、代码或图像形式输出,因此需第二个模型来将其与观察到的屏幕进行判断。我们认为用于验证的世界模型应在观察值编码的空间中进行预测,提出LGWM,这是一种无解码器、动作条件的世界模型,直接预测下一屏幕的表示,在185万真实GUI过渡上训练且无需语义标注。验证简化为向量比较,相同信号可揭示不匹配是否有害。我们在RSWT-BENCH上进行评估,该基准中每个凭证屏幕均在合法和劫持过渡下出现,因此仅看屏幕的检测器按构造处于随机水平。该无训练分数达到0.987的AUC,每次决策耗时17毫秒,与最强的闭源视觉语言模型(VLM)相当,且比生成式GUI世界模型的AUC高约10个点,延迟低三个数量级以上。残差方向在区分有害与良性违规时达到0.953的AUC,而提示式VLM接近随机。进一步分析表明,该预测是可用的未来状态而非异常分数。世界模型大多用作模拟器或规划器;我们的结果指出了其第三个角色——验证,在该角色中,在表示空间进行预测是自然设计。
英文摘要:
A login screen that appears after a tap on Sign in is expected; the same screen after a tap on View order is an attack. For GUI agents, safety is therefore a property of the transition rather than of the screen, and a monitor that inspects only screens can be defeated by reusing a legitimate one. Judging a transition requires an expectation of what should have followed the action. Existing GUI world models provide one, but they output it as text, code, or images, so checking it against the observed screen requires a second model to judge the two. We argue that a world model meant for verification should instead predict in the space in which observations are encoded, and present LGWM, a decoder-free, action-conditioned world model that predicts the representation of the next screen directly, trained without semantic annotation on 1.85M real GUI transitions. Verification reduces to a vector comparison, and the same signal reveals whether a mismatch is harmful. We evaluate on RSWT-BENCH, a diagnostic where each credential screen appears under both a legitimate and a hijacked transition, so detectors that see only the screen are at chance by construction. The training-free score reaches 0.987 AUC at 17 ms per decision, on par with the strongest closed-source VLMs and about ten AUC points above generative GUI world models at over three orders of magnitude lower latency. The residual direction reaches 0.953 AUC at separating harmful from benign violations, where prompted VLMs are near chance. Further analyses show that the prediction is a usable future state rather than an anomaly score. World models have mostly served as simulators or planners; our results point to a third role, verification, for which predicting in representation space is the natural design.