arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

使用小型语言模型逆向工程机器学习流水线结构

Using Small Language Models to Reverse-Engineer Machine Learning Pipelines Structures

Nicolas Lacroix, Frederic Precioso, Mireille Blay-Fornarino, Sebastien Mosser

arXiv 2610.10261首次发表:更新:

发表机构

Inria; CNRS; I3S; McMaster University(法国国家信息与自动化研究所; 法国国家科学研究中心; I3S研究所; 麦克马斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估小型语言模型(SLMs)在从源代码中提取机器学习流水线阶段的能力,发现分类法措辞显著影响性能,SLM表现良好但未超越现有分类器,且不同方法导致对ML实践的不同见解,揭示现有分类方法的局限性。

AI 中文摘要

背景:一旦定义了构成机器学习(ML)流水线的阶段分类法(例如,数据预处理、建模等),从源代码中提取这些阶段对于更好地理解ML实践至关重要。然而,由ML的持续演变(例如,算法、数据集)所导致的多样性使得这一任务具有挑战性。现有方法要么依赖不可扩展的人工标注,要么依赖不能适当支持领域多样性的分类器。这些局限性要求更可靠的解决方案。目标:我们评估小型语言模型(SLMs)是否可以利用其代码理解和分类能力来解决这些局限性,并增强我们对ML实践的理解。方法:我们基于代表当前最先进技术局限性的两项相关参考工作进行了验证性研究。我们首先使用Cochran's Q检验比较了多个SLM,然后通过两次McNemar检验将最佳性能模型与参考研究进行评估。额外的Cochran's Q检验考察了分类法定义变化如何影响SLM性能。最后,拟合优度检验比较了SLM分类得出的ML实践见解与先前研究得出的见解。结果:首先,我们发现分类法措辞显著影响分类性能。其次,性能最佳的SLM取得了良好结果,但并未超越其他分类器。第三,在探索数据科学家的实践时,三种分类方法导致了显著不同的见解,且效应量各异。结论:现有分类方法的局限性偏倚了我们对ML实践的理解。虽然当前的SLM在无需预先微调的情况下显示出有前景的结果,但它们仍然表现出常见的局限性,此外,推理成本高昂挑战了它们在大规模研究中的适用性。

英文摘要

Context: Once defined a taxonomy of stages structuring Machine Learning (ML) pipelines (e.g. Data Preprocessing, Modeling...), extracting these stages from source code is key for better understanding ML practices. However, the diversity caused by the constant evolution of ML (e.g., algorithms, datasets) makes this task challenging. Existing approaches either rely on non-scalable manual labeling or on classifiers that do not properly support domain's diversity. These limitations call for more reliable solutions. Objective: We evaluate whether Small Language Models (SLMs) can leverage their code understanding and classification abilities to address these limitations, and enhance our understanding of practices in ML. Method: We conduct a confirmatory study based on two relevant reference works representing current limitations in the state-of-the-art. We first compare several SLMs using Cochran's Q test, then evaluate the best-performing model against reference studies via two McNemar's tests. An additional Cochran's Q test examines how taxonomy definition variations affect the SLM performance. Finally, goodness-of-fit tests compare ML practice insights from SLM classification with those from prior studies. Results: First, we found that the taxonomy wording significantly impacts classification performance. Second, the best performing SLM yielded good results, yet, without outperforming other classifiers. Third, the three classification methods led to significantly different insights, with varying effect sizes, when exploring practices of data scientists. Conclusions: Limitations of existing classification methods bias our understanding of ML practices. While current SLMs show promising results without prior fine-tuning, they still exhibit common limitations, in addition to inference high costs challenging their applicability in large-scale studies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑