arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

图形用户界面代理相信它们的眼睛吗?诊断状态信念对像素与结构的依赖

Do GUI Agents Believe Their Eyes? Diagnosing State-Belief Reliance on Pixels versus Structure

Guijia Zhang, Yuxun Chen, Yuheng Qi, Harry Yang

arXiv 2607.04334首次发表:更新:

发表机构

Shenzhen University; The Hong Kong University of Science and Technology(深圳大学; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多模态GUI代理状态信念来源,通过对310个真实探针进行单通道干预形式化视觉状态依赖并测量,核心指标是感知融合差距,发现文本状态信念依赖结构,图像精度高,错误会导致行动失败。

AI 中文摘要

多模态GUI代理通过屏幕截图的渲染像素和序列化结构(如DOM或可访问性树)读取界面。现有基准测试不关注代理状态信念是否来自像素。我们形式化视觉状态依赖,通过对310个真实网络、移动和桌面探针的单通道干预进行测量。核心指标是感知融合差距,结果显示文本状态信念依赖结构,错误会导致行动失败。

英文摘要

Multimodal GUI agents read an interface through two redundant channels: the rendered pixels of a screenshot and a serialized structure such as a document object model or accessibility tree. Before acting, an agent forms a belief about the current interface state, but existing benchmarks score task success, element grounding, or attack resistance and do not ask whether that belief is drawn from the pixels. We formalize visual state reliance, the attribution of a state belief to pixels, structure, or priors, and measure it with paired single-channel interventions over 735 probes spanning real web, mobile, and desktop interfaces, of which 225 are zero-edit divergences mined from live production websites, all scored by deterministic forced choice with no model judge. Our central metric is the Perception-Fusion Gap (PFG), the fraction of probes a model perceives correctly yet resolves toward structure under conflict; a stricter variant that re-verifies perception on a tight crop of the target region leaves the gap intact. Across models from four vendors, textual state beliefs defer to structure while image-only accuracy stays near ceiling, and on unedited stale snapshots from live pages the same models follow the outdated structure on up to 0.88 of probes. A white-box ablation traces the textual effect to a single copied structural value, and gradient attribution shows the visual evidence is processed yet overridden. In live multi-step environments, one mis-sourced belief at the first step compounds into task failure with a self-recovery rate of at most 0.03. Comparing four mitigations on identical probes, prompt-level cues fail at the action level, certificate checks buy safety with refusals, and a training-free consistency gate is alone in reducing both hijack and task error. Visual state reliance thus gives a measurable diagnostic of whether agent state beliefs are visually grounded.

Comments17 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑