arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16022cs.SE

OpenHarmony Bench:评估LLM与编码智能体在OpenHarmony应用开发中的表现

OpenHarmony Bench: Evaluating LLMs and Coding Agents on OpenHarmony App Development

Li Li, Han Hu, Tianjian Zhang, Xin Peng, Fangzhu Mao, Qingyu Zhang, Xiaoheng Xie, Zhongmin Tang, Zhihao Lin, Haolin Ruan, Miaomiao Dong, Liuchuan Zhu, Yue Li, C… 展开作者

Li Li, Han Hu, Tianjian Zhang, Xin Peng, Fangzhu Mao, Qingyu Zhang, Xiaoheng Xie, Zhongmin Tang, Zhihao Lin, Haolin Ruan, Miaomiao Dong, Liuchuan Zhu, Yue Li, Chi Chen, Wenkang Zhong, Mingfei Zhang, Yang Yu, Bo Sun, Chaorui Zhang, Weixi Zhang, Wei Han, Bo Bai, Kui Liu, Gang Fan, Siru Liu, Jiaqian Zhou, Jiali Sun, Yunbiao Dong, Wenhao Zhong, Yunhong Xu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出OPENHARMONY BENCH应用级编码基准,评估8个LLM驱动的DevEco Code在三类OpenHarmony ArkTS应用任务中的表现,发现新一代模型任务完成率更高、构建性接近饱和但行为正确性不足、spec驱动任务完成率最低。

中文摘要 AI 辅助

我们推出OPENHARMONY BENCH,这是一个用于评估基于大语言模型(LLM)的编码智能体在OpenHarmony ArkTS应用开发能力的应用级编码基准。与函数级基准不同,它评估完整的应用级变更:每个任务要求智能体修改一个可构建的ArkTS项目,以实现请求的端到端行为,涉及UI状态、数据持久化、构建配置和平台API。该基准会在设备上安装并运行交付的应用,以检查行为是否可观测。它涵盖三类输入源:自然语言功能请求(new-feature)、结构化场景规范(spec-driven)和bug描述(bug-fix)。该基准包含153个顶级任务和242个功能点(F-points,即一个可执行的行为检查项),快照版本包含32个new-feature任务、50个spec-driven任务(对应139个F-points)以及71个bug-fix任务。主排行榜基于顶级任务而非独立加权的F-points进行评分。我们描述了基准的构建过程、统计数据以及构建-测试评估流程,并针对每种配置开展三次独立的全套件运行,用8个LLM评估了DevEco Code。得出三项发现:其一,在评估的模型家族对中,新一代模型比前代完成更多任务;其二,可构建性接近饱和,而行为正确性尚未达到:平均最终构建成功率为94.77%至100.00%,而平均任务完成率为48.36%至58.39%;其三,在全检查任务评分规则下,spec-driven任务的任务完成率最低,无任何配置超过35%。代码、数据、任务、参考解决方案、测试、评估脚本和排行榜均通过官方OPENHARMONY BENCH网站发布,网址为this https URL。

英文摘要

We present OPENHARMONY BENCH, an app-level coding benchmark for evaluating LLM-based coding agents on OpenHarmony ArkTS applications. Unlike function-level benchmarks, it evaluates complete app-level changes: each task requires an agent to modify a buildable ArkTS project so that a requested behavior works end to end, involving UI state, data persistence, build configuration, and platform APIs. The benchmark installs and drives the delivered application on a device to check whether the behavior is observable. It covers three input sources: natural-language feature requests (new-feature), structured scenario specifications (spec-driven), and bug descriptions (bug-fix). The benchmark contains 153 top-level tasks and 242 Feature points (F-points), where an F-point is one executable behavior check. The snapshot includes 32 new-feature tasks, 50 spec-driven tasks with 139 F-points, and 71 bug-fix tasks. The main leaderboard is scored over top-level tasks rather than independently weighted F-points. We describe the benchmark construction, statistics, and build-and-test evaluation pipeline, and evaluate DevEco Code with eight LLMs across three independent full-suite runs per configuration. Three findings emerge. First, newer generations complete more tasks than their predecessors within evaluated model-family pairs. Second, buildability is close to saturated while behavioral correctness is not: mean Final Build Success Rate is 94.77% to 100.00%, whereas mean Task Completion is 48.36% to 58.39%. Third, spec-driven tasks have the lowest Task Completion under all-checks task scoring, with no configuration exceeding 35%. The code, data, tasks, reference solutions, tests, evaluation scripts, and leaderboard are released through the official OPENHARMONY BENCH website at https://bench.matrix.openharmony.cn/.

↑