发表机构
Uber Technologies, Inc.(优步科技公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DragonCrawl是基于GPT-4o多模态的移动端到端测试框架,通过验证用户流程阻止问题提交,缩短测试入门时间,节省大量维护工作量,实现规模化持续质量保证。
AI 中文摘要
随着移动应用复杂度提升,传统端到端(E2E)测试框架面临UI波动、维护开销大、跨平台可扩展性差等问题。本文提出DragonCrawl,一种用于持续回归测试的AI驱动移动测试系统,其已从基于嵌入的相似度匹配演进为利用大语言模型的生成式意图推理。与此前聚焦探索性测试和崩溃检测的LLM测试研究不同,DragonCrawl在每次代码变更时验证特定用户流程,阻止破坏关键功能的提交。通过利用GPT-4o的多模态能力,DragonCrawl在CI/CD流水线中持续运行的1013项自动化测试,在iOS上达到91.6%的通过率,在Android上达到92.2%。该系统将测试入门时间从96-120小时缩短至4小时以内,已节省约27个开发者年的测试维护工作量。本文阐述了从V1(语义嵌入匹配)到V2(生成式意图推理)的架构演进,讨论了包括token爆炸和内存限制在内的实现挑战,并报告了生产部署的运营经验。多模态视觉用于最终状态检测、工具调用用于后端状态转换的集成,使回归测试能全面连接UI交互与系统状态。研究结果表明,AI驱动的测试可在保持稳定性的同时消除传统自动化测试的脆弱性,实现规模化的持续质量保证。
英文摘要
As mobile applications grow in complexity, traditional End-to-End (E2E) testing frameworks struggle with UI volatility, maintenance overhead, and cross-platform scalability. This paper presents DragonCrawl, an AI-driven mobile testing system for continuous regression testing that has evolved from embedding-based similarity matching to generative intent-based reasoning using large language models. Unlike prior LLM-based testing research focused on exploratory testing and crash detection, DragonCrawl validates specific user flows on every code change, blocking commits that break critical functionality. By leveraging GPT-4o's multimodal capabilities, DragonCrawl achieves 91.6% pass rate on iOS and 92.2% on Android across 1,013 automated tests running continuously in CI/CD pipelines. The system reduces test onboarding time from 96-120 hours to under 4 hours and has saved an estimated 27 developer years in test maintenance effort. We present the architectural evolution from V1 (semantic embedding matching) to V2 (generative intent-based reasoning), discuss implementation challenges including token explosion and memory constraints, and report operational experience from production deployment. The integration of multimodal vision for end-state detection and tool calling for backend state transitions enables comprehensive regression testing that bridges UI interactions with system state. Our results demonstrate that AI-driven testing can maintain stability while eliminating the brittleness of traditional automated tests, enabling continuous quality assurance at scale.
Comments12 pages, 6 figures, 6 pages