arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于自动化跨领域机器学习任务类型识别的大型语言模型:基准数据集与评估

Large Language Models for Automated Cross-Domain Machine Learning Task Type Identification: A Benchmark Dataset and Evaluation

Petros Tsialis, Steffen Limmer, Tobias Rodemann, Martin Heckmann

arXiv 2609.35335首次发表:更新:

发表机构

University of Applied Sciences Aalen; Honda Research Institute Europe(阿伦应用科学大学; 本田欧洲研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出利用大型语言模型自动识别机器学习任务类型,并发布含625个数据集的基准,实验表明该方法在表格和跨领域场景中优于AutoGluon,但小型模型存在精度与部署的权衡。

AI 中文摘要

机器学习任务类型识别对于构建有效的机器学习流水线至关重要,然而在实践中,该识别通常由人工指定。我们研究了大型语言模型(LLM)是否能够在用户仅提供目标特征的情况下,直接从数据集层面的信息推断出数据领域和下游预测任务。与基于LLM的系统一同,我们还发布了一个包含625个公开表格和时间序列数据集的带注释基准。我们在三种设置下评估了所提出的方法:(i)与成熟的AutoML启发式方法相比的表格数据集,(ii)跨表格和时间序列数据集的跨领域评估,以及(iii)使用较小本地模型的实际部署场景。结果显示,基于LLM的任务类型识别具有一致的优势,但在异构和资源受限的设置中难度增加。基于LLM的方法在表格设置中优于AutoGluon,达到了0.98的宏F1分数,而AutoGluon为0.93。在跨领域设置中,最佳模型达到了0.90的宏F1分数,而较小的可本地部署模型达到了0.75,这表明了部署可行性与准确性之间的权衡。

英文摘要

Machine learning task type identification is essential for constructing valid ML pipelines, yet in practice it is typically specified manually. We investigate whether large language models (LLMs) can infer both the data domain and the downstream prediction task directly from dataset-level information when only the target feature is provided by the user. Together with our LLM-based system we also release an annotated benchmark comprising 625 public tabular and time series datasets. We evaluate the proposed approach in three settings: (i) tabular datasets in comparison with established AutoML heuristics, (ii) cross-domain evaluation across tabular and time series datasets, and (iii) a practical deployment scenario using smaller local models. The results show consistent advantages for LLM-based task type identification, with increasing difficulty in heterogeneous and resource-constrained settings. LLM-based approaches outperform AutoGluon in the tabular setting, reaching 0.98 F1 macro compared to 0.93. In the cross-domain setting, the best model achieves 0.90 F1 macro, while smaller locally deployable models reach 0.75, indicating a trade-off between deployment feasibility and accuracy.

Comments25 pages

Journal refInternational Conference on Automated Machine Learning (AutoML 26), September 2026, Ljubljana, Slovenia

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑