OpenHarmony Bench:评估LLM与编码智能体在OpenHarmony应用开发中的表现
OpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development
浏览论文内容
中文总结 AI 辅助
该研究推出OPENHARMONY BENCH应用级编码基准,评估8个LLM驱动的DevEco Code在三类OpenHarmony ArkTS应用任务中的表现,发现新一代模型任务完成率更高、构建性接近饱和但行为正确性不足、spec驱动任务完成率最低。
中文摘要 AI 辅助
我们推出OPENHARMONY BENCH,这是一个用于评估基于大语言模型(LLM)的编码智能体在OpenHarmony ArkTS应用开发能力的应用级编码基准。与函数级基准不同,它评估完整的应用级变更:每个任务要求智能体修改一个可构建的ArkTS项目,以实现请求的端到端行为,涉及UI状态、数据持久化、构建配置和平台API。该基准会在设备上安装并运行交付的应用,以检查行为是否可观测。它涵盖三类输入源:自然语言功能请求(new-feature)、结构化场景规范(spec-driven)和bug描述(bug-fix)。该基准包含153个顶级任务和242个功能点(F-points,即一个可执行的行为检查项),快照版本包含32个new-feature任务、50个spec-driven任务(对应139个F-points)以及71个bug-fix任务。主排行榜基于顶级任务而非独立加权的F-points进行评分。我们描述了基准的构建过程、统计数据以及构建-测试评估流程,并针对每种配置开展三次独立的全套件运行,用8个LLM评估了DevEco Code。得出三项发现:其一,在评估的模型家族对中,新一代模型比前代完成更多任务;其二,可构建性接近饱和,而行为正确性尚未达到:平均最终构建成功率为94.77%至100.00%,而平均任务完成率为48.36%至58.39%;其三,在全检查任务评分规则下,spec-driven任务的任务完成率最低,无任何配置超过35%。代码、数据、任务、参考解决方案、测试、评估脚本和排行榜均通过官方OPENHARMONY BENCH网站发布,网址为this https URL。
英文摘要
We present OPENHARMONY BENCH, an app-level coding benchmark for evaluating LLM-based coding agents on OpenHarmony ArkTS applications. Unlike function-level benchmarks, it evaluates complete app-level changes: each task requires an agent to modify a buildable ArkTS project so that a requested behavior works end to end, involving UI state, data persistence, build configuration, and platform APIs. The benchmark installs and drives the delivered application on a device to check whether the behavior is observable. It covers three input sources: natural-language feature requests (new-feature), structured scenario specifications (spec-driven), and bug descriptions (bug-fix). The benchmark contains 153 top-level tasks and 242 Feature points (F-points), where an F-point is one executable behavior check. The snapshot includes 32 new-feature tasks, 50 spec-driven tasks with 139 F-points, and 71 bug-fix tasks. The main leaderboard is scored over top-level tasks rather than independently weighted F-points. We describe the benchmark construction, statistics, and build-and-test evaluation pipeline, and evaluate DevEco Code with eight LLMs across three independent full-suite runs per configuration. Three findings emerge. First, newer generations complete more tasks than their predecessors within evaluated model-family pairs. Second, buildability is close to saturated while behavioral correctness is not: mean Final Build Success Rate is 94.77% to 100.00%, whereas mean Task Completion is 48.36% to 58.39%. Third, spec-driven tasks have the lowest Task Completion under all-checks task scoring, with no configuration exceeding 35%. The code, data, tasks, reference solutions, tests, evaluation scripts, and leaderboard are released through the official OPENHARMONY BENCH website at https://bench.matrix.openharmony.cn/.