AI 中文总结
本研究提出AppEval基准框架,评估基于大语言模型的移动应用修复智能体在ArkTS、Swift、Kotlin平台的表现,发现其性能因智能体而异,且运行时感知的接受标准对有意义的比较至关重要。
AI 中文摘要
仓库级大语言模型智能体的评估通常基于测试在构建主机上运行的项目,但目前仍不清楚它们的修复是否能通过移动应用的构建-安装-启动-测试边界,在该边界中,缺失的软件开发工具包(SDK)、离线设备或断言前崩溃可能被误认为是程序故障。本文提出AppEval,这是一个针对HarmonyOS/ArkTS、iOS/Swift及Android/Kotlin移动应用修复的基准和原生工具链评估框架。每个任务将隐藏的行为测试与参考生产修复分开,仅当相同的已安装应用目标在缺陷版本上达到断言失败且在修复后通过测试时,任务才被接受;基础设施故障则视为不同的结果。一种通用模式将该契约映射到每个平台的构建系统、运行时和测试运行器。经过审核的Android分区包含来自24个可独立构建仓库的200个已接受的插装任务。在这些任务上,五个智能体的Pass@1(前1个通过的成功率)在22.00%至90.50%之间,在相同动态预言下的差距为68.50个百分点。这些结果表明,移动应用修复性能强烈依赖于所评估的智能体,同时也证明了基于运行时感知的接受标准对于有意义的比较是必要的。本文的定量发现仅针对Android;在得出跨平台泛化结论之前,还需要经过审核的iOS和HarmonyOS结果。
英文摘要
Repository-level LLM agents are typically evaluated on projects whose tests run on the build host. It remains unclear whether their repairs survive the mobile build-install-launch-test boundary, where a missing SDK, offline device, or pre-assertion crash can be mistaken for a program failure. We present AppEval, a benchmark and native-toolchain evaluation framework for mobile application repair across HarmonyOS/ArkTS, iOS/Swift, and Android/Kotlin. Each task separates a hidden behavior test from the reference production fix and is accepted only when the same installed-app target reaches an assertion failure on the defective revision and passes after the fix; infrastructure failures remain a distinct outcome. A common schema maps this contract to each platform's build system, runtime, and test runner. The audited Android partition contains 200 accepted instrumentation tasks from 24 independently buildable repositories. On these tasks, five agents achieve Pass@1 between 22.00% and 90.50%, a 68.50-percentage-point spread under the same dynamic oracle. These results show that mobile repair performance depends strongly on the evaluated agent while demonstrating why runtime-aware acceptance is necessary for meaningful comparison. The quantitative findings in this paper are Android-specific; audited iOS and HarmonyOS results are required before drawing cross-platform generalization conclusions.
Comments9 pages, 2 figures, 5 tables