发表机构
Delft University of Technology(代尔夫特理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过挖掘 GitHub 和 Kaggle 上的 Jupyter notebook,提取并分类了超过百万条反馈语句,构建了分类体系,揭示了探索性反馈的主导地位及对静默失败的防御实践。
AI 中文摘要
Jupyter notebook 中的机器学习开发是迭代且由反馈驱动的。从业者编写语句来揭示程序执行的信息,并利用这些信息决定下一步行动。我们将这些语句称为反馈语句,并识别出两种形式:用于可视化检查而显示值的探索性语句,以及通过断言以编程方式强制执行条件的验证性语句。许多机器学习失败不会以异常形式显现,因此逃过了主导先前机器学习 notebook 研究的基于崩溃的分析。本研究通过表征编码了从业者对代码应做什么以及可能出错的思维模型的反馈语句,来考察从业者检查什么以捕获否则会静默通过的失败。我们从 GitHub 和 Kaggle 挖掘了 297,851 个公开的 Python Jupyter notebook,并提取了 1,092,780 条反馈语句。我们通过从 CodeBERT 嵌入获得的语义簇中进行比例分层抽样,抽取了 816 条语句,并应用扎根理论和开放编码对每条语句进行标记和分析。我们贡献了一个机器学习 notebook 中反馈语句的分类体系,该体系按语句的功能意图和其出现的机器学习流水线阶段进行组织。该分类体系揭示,反馈绝大多数是探索性的,并且这两个平台承载了性质不同的机器学习工作模式。将我们的分类体系映射到现有的崩溃分类体系表明,它捕获了崩溃分析无法观察到的针对静默失败的防御性实践。我们的发现表明,在机器学习开发者实践的研究中,notebook 源码应被视为一个混杂因素,为 notebook 工具提供了机会,并激发了对静默机器学习失败的实证研究。我们发布了包含 1,092,780 条反馈语句的语料库和编码手册,以支持复现和工具研究。
英文摘要
Machine learning development in Jupyter notebooks is iterative and feedback-driven. Practitioners author statements that reveal information about program execution and use it to decide what to do next. We call these feedback statements and identify two forms: exploratory statements that display values for visual inspection, and validation statements that enforce conditions programmatically through assertions. Many ML failures do not surface as exceptions and thus escape the crash-based analyses that dominate prior work on ML notebooks. This study examines what practitioners check to catch failures that would otherwise pass silently, by characterizing feedback statements that encode the practitioner's mental model of what the code should do and what could go wrong. We mine 297,851 public Python Jupyter notebooks from GitHub and Kaggle and extract 1,092,780 feedback statements. We sample 816 statements through proportional stratified sampling from semantic clusters obtained from CodeBERT embeddings, and apply grounded theory and open coding to label and analyze each one. We contribute a taxonomy of feedback statements in ML notebooks, organized along the functional intent of the statement and the ML pipeline stage in which it appears. The taxonomy reveals that feedback is overwhelmingly exploratory, and that the two platforms host qualitatively different modes of ML work. Mapping our taxonomy to an existing crash taxonomy shows that it captures defensive practices against silent failures that crash analysis cannot observe. Our findings indicate that notebook source should be treated as a confounder in studies of ML developer practice, surface opportunities for notebook tooling, and motivate empirical study of silent ML failures. We release the corpus of 1,092,780 feedback statements and the codebook to support replication and tooling research.