发表机构
Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory(上海交通大学; 上海人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对GUI智能体,提出基于层次化运动(连贯、结构、生物)的验证码框架MVCAP及基准MVCAP-Bench,人类准确率99.6%而最佳智能体仅16.8%,揭示动态背景伪装造成显著人机感知差距。
AI 中文摘要
现有的大多数视觉验证码在空间上仍然可解:所需信息通过静态外观、局部结构和界面状态暴露。多模态大语言模型(MLLMs)和图形用户界面(GUI)智能体的进步削弱了这一假设,这些智能体展现出强大的视觉感知、推理和浏览器交互能力。我们提出了运动视觉验证码(MVCAP),一种基于层次化运动的验证码框架,其中目标语义被实例化为运动定义的前景结构,并且只能通过从动态演化的背景中进行时间分离来恢复。基于这一共同原则,MVCAP在三个感知递进层级上实例化:连贯运动、结构运动和生物运动。为评估该框架,我们引入了MVCAP-Bench,一个基于浏览器的基准测试,包含600个实时验证码实例,以及一个匹配的前景-only控制基准MVCAP-Bench-FG。我们评估了人类、Browser Use智能体、原生计算机使用智能体,以及从相同实例派生的补充离线VQA设置。结果显示存在显著的人类-智能体差距:在完整MVCAP-Bench上,人类准确率达到99.6%,而最佳GUI智能体仅达到16.8%,接近六选一的随机水平。前景-only控制进一步表明,关键困难来自动态背景伪装,而非答案格式或浏览器交互本身。这些发现识别出可测量的人类-智能体感知差距,并将MVCAP-Bench定位为研究当前智能体中运动定义感知的基准。
英文摘要
Most existing visual CAPTCHAs remain spatially solvable: the required information is exposed by static appearance, local structure, and interface state. This assumption is weakened by advances in multimodal large language models (MLLMs) and Graphical User Interface (GUI) agents, which exhibit strong visual perception, reasoning, and browser interaction capabilities. We propose Motion Vision CAPTCHA (MVCAP), a hierarchical motion-based CAPTCHA framework in which target semantics are instantiated as motion-defined foreground structures and become recoverable only through temporal segregation from a dynamically evolving background. Built on this shared principle, MVCAP is instantiated in three perceptually progressive levels: coherent motion, structural motion, and biological motion. To evaluate this framework, we introduce MVCAP-Bench, a browser-based benchmark with 600 live CAPTCHA instances, together with a matched foreground-only control benchmark, MVCAP-Bench-FG. We evaluate humans, Browser Use agents, native computer use agents, and a supplementary offline VQA setting derived from the same instances. Results reveal a substantial human--agent gap: on the full MVCAP-Bench, human accuracy reaches 99.6%, whereas the best GUI agent achieves only 16.8%, close to the six-way chance level. The foreground-only control further shows that the key difficulty comes from dynamic background camouflage rather than answer format or browser interaction alone. These findings identify a measurable human--agent perception gap and position MVCAP-Bench as a benchmark for studying motion-defined perception in current agents.
CommentsAccepted at ACM Multimedia 2026. 10 pages, 5 figures