arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越排行榜:密集模型与混合专家模型在自动程序修复中的多维评估

Beyond the Leaderboard: Multi-Dimensional Evaluation of Dense and Mixture-of-Experts Models for Automated Program Repair

Anvi Kalpesh Shah, Umamaheswara Sharma B

arXiv 2610.08173首次发表:更新:

发表机构

National Institute of Technology, Calicut(卡利卡特国家理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自动程序修复,提出加权质量指数(QI)综合评估功能正确性、可维护性、安全性和效率,发现MoE模型以更少激活参数达到与较大密集模型相当的正确性,凸显激活参数作为评估视角的重要性。

AI 中文摘要

使用语言模型进行自动程序修复(APR)通常通过生成的补丁是否通过测试套件来评估,这可能会掩盖可维护性、安全性和计算成本方面的差异。我们提出了一种加权质量指数(QI),其灵感来源于ISO/IEC 25010软件质量模型,该指数在可配置的加权方案下结合了功能正确性、可维护性、安全性和生成效率。我们在40个QuixBugs和90个Defects4J缺陷上评估了三个密集Qwen2.5-Coder模型(3B、7B、14B)以及16B参数的DeepSeek-Coder-V2-Lite混合专家(MoE)模型(2.4B激活参数),所有模型均在相同的本地硬件上运行,以控制基础设施效应。模型排名随加权方案的变化而变化,表明单一指标评估可能隐藏权衡。MoE模型与7B和14B密集模型在正确性上几乎没有统计学显著差异(McNemar精确检验),同时使用的激活参数少3-6倍,而三个密集规模之间的正确性显著增加。这些结果表明,对于稀疏代码模型,激活参数数量可能比总参数数量更具信息性。

英文摘要

Automated Program Repair (APR) with language models is usually evaluated by whether a generated patch passes the test suite, which can hide differences in maintainability, security, and computational cost. We propose a Weighted Quality Index (QI), inspired by the ISO/IEC 25010 software quality model, that combines functional correctness, maintainability, security, and generation efficiency under configurable weighting schemes. We evaluate three dense Qwen2.5-Coder models (3B, 7B, 14B) and the 16B-parameter DeepSeek-Coder-V2-Lite Mixture-of-Experts (MoE) model (2.4B active parameters) on 40 QuixBugs and 90 Defects4J bugs, all run locally on identical hardware to control for infrastructure effects. Model rankings change with the weighting scheme, showing that single-metric evaluation can hide trade-offs. The MoE model shows almost no statistically significant difference in correctness from the 7B and 14B dense models (McNemar's exact test) while using 3-6 times fewer active parameters, whereas correctness increases significantly across the three dense scales. These results suggest that active parameter count can be a more informative lens than total parameter count for sparse code models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑