arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Web开发中代码驱动智能体测试的框架与基准

Framework and Benchmark for Code-Driven Agentic Testing in Web Development

Bin Hong, Zhenchao Zhang, Jiyuan He, Kai Zhang, Zhenya Huang

arXiv 2609.00081首次发表:更新:

AI 中文总结

本研究针对现有GUI测试评估的局限,提出代码驱动智能体测试(CAT)范式及配套框架与基准,实验发现主流VLM的漏洞发现能力远未满足实际Web开发测试需求。

AI 中文摘要

端到端GUI测试是验证Web应用的关键,但现有评估依赖预定义检查表,且局限于Web生成基准的数据和框架,导致视觉语言模型(VLM)的漏洞发现能力未得到系统测试。我们提出代码驱动智能体测试(CAT)范式,该范式中智能体编写Playwright代码驱动浏览器、收集反馈并自主探索Web应用以发现漏洞。我们通过CATJudge实例化CAT,这是一个将Browser-Use和Computer-Use工具统一在单一环境中的智能体框架;还构建了CATTest,这是一个包含102个AI生成Web应用的基准,这些应用由人机密切协作打造,具有复杂交互和细微缺陷。对主流VLM的实验显示,所有评估模型表现均不佳,揭示了当前VLM能力与AI Web开发中实际测试需求之间存在明显差距。我们在该httpsURL发布代码和数据。

英文摘要

End-to-end GUI testing is essential for verifying web applications, yet existing evaluations rely on predefined checklists and are confined to the data and frameworks of web generation benchmarks, leaving the bug-discovery ability of vision-language models (VLMs) systematically untested. We introduce \textbf{C}ode-driven \textbf{A}gentic \textbf{T}esting (CAT), a paradigm in which the agent writes Playwright code to drive the browser, gathers feedback, and autonomously explores web applications to uncover bugs. We instantiate CAT with CATJudge, an agentic framework that unifies Browser-Use and Computer-Use tools within a single environment and CATTest, a benchmark of 102 AI-generated web applications with carefully annotated bugs, built through close human-AI collaboration to feature complex interactions and subtle defects. Experiments with mainstream VLMs show that all evaluated models perform poorly, revealing a clear gap between current VLM capabilities and the demands of real-world testing in AI web development. We release our code and data at https://github.com/SleepyWithoutCoffee/CATJudge.

Comments30 pages, 20 figures, EMNLP'26 Main Conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑