arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10610cs.SEcs.LG

代码质量与机器学习性能之间的关系:一项大规模实证研究

On the Relation between Code Quality and Machine Learning Performance: A Large-scale Empirical Study

Marius Mignard, Steven Costiou, Anne Etien

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过分析Kaggle上265,363个Python笔记本,发现通用代码质量与机器学习性能无关,而机器学习特定违规与性能呈小负相关,流行度和代码专业知识不指示质量或性能,竞赛专业知识则与更好性能相关。

中文摘要 AI 辅助

背景:计算笔记本是机器学习(ML)开发的标准环境。在机器学习社区中,模型性能通常被视为首要指标,而代码质量则被视为次要问题。这种优先级排序依赖于一个在很大程度上未经检验的假设,即代码质量与机器学习性能无关。实践者还会重用现有代码,这些代码可能来自通过社会信号(流行度、作者专业知识)选择的笔记本,而这些信号作为质量代理的可靠性从未被评估。目标:我们实证研究了笔记本中代码质量与机器学习性能之间的关系,并评估了流行度和作者专业知识是否能为代码质量或性能提供指示。方法:我们对提交到Kaggle竞赛的265,363个Python笔记本进行了一项大规模实证研究。我们使用两个静态分析工具评估代码质量:Pylint,用于捕获通用Python代码质量;以及SonarQube,配置了包含34条针对数据科学和机器学习特定实践的规则的分析文件。结果:代码质量与性能之间的关系取决于所考虑的质量概念。通用Python代码质量与机器学习性能脱钩,在所有观察中显示出可忽略或非显著的关联。相比之下,机器学习特定的违规行为与性能之间存在一致的、小的负关联,且在所有观察中持续存在。笔记本的流行度不提供关于代码质量或性能的信息。代码专业知识不提供关于质量或性能的信息,但竞赛专业知识与更好的性能、更少的机器学习特定违规行为以及略多的Python错误和重构违规行为相关。

英文摘要

Context: Computational notebooks are the standard environment for machine learning (ML) development. Within the ML community, model performance is often the primary considered metric, and code quality is treated as a secondary concern. This prioritization relies on a largely untested assumption that code quality and ML performance are unrelated. Practitioners also reuse existing code that may come from notebooks selected through social signals (popularity, author expertise) whose reliability as quality proxies has never been assessed. Objective: We empirically investigated the relationship between code quality and ML performance in notebooks, and evaluated whether popularity and author expertise give indication on code quality or performance. Method: We conducted a large-scale empirical study of 265,363 Python notebooks submitted to Kaggle competitions. We assessed code quality with two static analysis tools: Pylint, capturing general Python code quality, and SonarQube, configured with a profile of 34 rules targeting data-science and ML-specific practices. Results: The relationship between code quality and performance depends on the notion of quality considered. General Python code quality is decoupled from ML performance, showing negligible or non-significant correlations across all observations. In contrast, ML-specific violations exhibit a consistent, small negative association with performance that persists across all observations. The popularity of a notebook does not give information on the code quality or performance. Code expertise provides no information on quality or performance, but competition expertise correlates with better performance, fewer ML-specific violations, and slightly more Python errors and refactoring violations.

发表机构

  • Univ. Lille(里尔大学)
  • Inria(法国国家信息与自动化研究所)
  • CNRS(法国国家科学研究中心)
  • Centrale Lille(里尔中央理工学院)
  • UMR 9189 CRIStAL(UMR 9189 里尔计算机、信号与自动控制实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑