SAGE-Loop:具有试错修正与自适应集成的可靠闭环LLM驱动AutoML
SAGE-Loop: Reliable Closed-Loop LLM-Driven AutoML with Trial-and-Correction and Adaptive Ensembling
浏览论文内容
中文总结 AI 辅助
针对LLM驱动AutoML缺乏闭环修正与静态集成的问题,提出SAGE-Loop框架,通过多轮试错修复与自适应集成,在20个数据集上提升分类、回归和聚类任务的性能与稳定性。
中文摘要 AI 辅助
自动化机器学习(AutoML)正在重塑数据驱动的科学和工业实践,随着大型语言模型被引入AutoML,流水线可靠性变得与自动化效率同等重要。然而,现有的AutoML在执行过程中仍难以实现即时反馈和自适应优化,因此一旦运行偏离到次优或失败状态,就缺乏过程级修正机制。根本问题在于其单向流水线:中间失败通常被终止或绕过,而固定范式往往强化模型生成,却使集成决策保持静态,从而削弱了执行可靠性以及对结构多样性的受控使用。这表明,LLM驱动的AutoML需要具备试错-修正-改进的闭环能力,以及对模型多样性的基于证据的使用。为此,我们提出了SAGE-Loop,一个可靠的闭环、自适应、LLM驱动的AutoML框架,它在监督和无监督任务中执行多轮生成与验证以进行试错修复,并自适应选择集成策略,从而统一了模型的生成方式与使用方式。在20个公共数据集上,SAGE-Loop在分类、回归和聚类任务上持续提升了性能和稳定性。额外结果进一步表明其能够从执行失败中恢复并维持稳健的流水线行为。
英文摘要
Automated machine learning (AutoML) is reshaping data-driven science and industrial practice, and as large language models are introduced into AutoML, pipeline reliability becomes as important as automation efficiency. However, existing AutoML still struggles to realize instant feedback and adaptive optimization during execution, so once a run drifts into a suboptimal or failed state, it lacks a process-level correction mechanism. The fundamental pathology lies in its one-way pipeline: intermediate failures are typically terminated or bypassed, while fixed paradigms often strengthen model generation but leave ensemble decisions static, weakening both execution reliability and the controlled use of structural diversity. This indicates that LLM-driven AutoML needs a closed-loop ability for trial-correction-improvement together with evidence-based use of model diversity. To this end, we propose SAGE-Loop, a reliable closed-loop, self-adaptive, LLM-driven AutoML framework that performs multi-round generation and validation for trial-and-repair, and adaptively selects ensemble strategies in both supervised and unsupervised tasks, thereby unifying how to generate with how to use models. Across 20 public datasets, SAGE-Loop consistently improves performance and stability on classification, regression, and clustering tasks. Additional results further show its ability to recover from execution failures and maintain robust pipeline behavior.