发表机构
University of Waterloo; Google Inc.(滑铁卢大学; 谷歌公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究系统研究相关MLE任务洞察对智能体MLE系统的影响,提出MLE-InsightBench基准和MLE-InsightForge组装智能体,通过三种洞察类型提升性能,归一化私有测试分数提高123.9%。
AI 中文摘要
智能体机器学习工程(MLE)是AI4SE中一种新兴的应用,用于处理复杂的机器学习任务,也是迈向AI系统递归自我改进的一步。最近的智能体MLE系统展示了利用相关MLE任务洞察的价值:一些系统基于专家领域知识进行代码生成,这些知识是从同类MLE任务中隐式整理得到的;一些系统则通过循环解决MLE任务、收集记忆以惠及未来任务。然而,洞察的收集和注入仍然是临时性的,且现有评估往往不控制洞察来源,使得数据泄漏成为有效性的威胁。为弥补这些差距,我们系统地研究了相关MLE任务的洞察如何影响智能体MLE系统,以及其影响如何依赖于洞察类型和注入方法。我们引入了MLE-InsightBench,一个涵盖12个领域、包含160个Kaggle竞赛的基准。为防止泄漏,每个领域保留一个竞赛作为目标,其余148个竞赛仅当它们不可能受到目标任务影响时(例如通过后续版本或重用解决方案)才作为来源。我们还开发了MLE-InsightForge,一个洞察组装智能体,它探索来源竞赛、顶级人类解决方案和文章,以构建三种洞察类型:12个领域洞察总结最佳实践,148个竞赛洞察描述获胜解决方案,以及16个改进洞察捕捉通用的调试和优化策略。这些洞察以按需技能或前置手册的形式注入。在我们的评估中,提取的洞察相比无洞察基线,将归一化私有测试分数提高了123.9%。受控实验进一步表明,三种洞察类型是互补的;更好的注入模式(技能与手册)因任务而异,总体上技能更具成本效益;且这些收益在不同后端LLM上具有泛化性。
英文摘要
Agentic machine learning engineering (MLE) is an emerging AI4SE application for complex ML tasks and a step toward recursive self-improvement of AI systems. Recent agentic MLE systems show the value of leveraging insights from related MLE tasks: some systems condition code generation on expert domain knowledge, which is implicitly curated from peer MLE tasks; some systems have a loop of solving an MLE task, gathering memory to benefit future tasks. However, insight collection and injection remain ad hoc, and existing evaluations often do not control insight sources, making data leakage a threat to validity. To close these gaps, we systematically study how insights from related MLE tasks affect agentic MLE systems, and how their impact depends on insight type and injection method. We introduce MLE-InsightBench, a benchmark of 160 Kaggle competitions across 12 domains. To prevent leakage, one competition per domain is held out as the target, while the remaining 148 serve as sources only when they cannot have been influenced by the target task, such as through later editions or reused solutions. We also develop MLE-InsightForge, an insight assembly agent that explores source competitions, top human solutions, and writeups to construct three insight types: 12 domain insights summarizing best practices, 148 competition insights describing winning solutions, and 16 improvement insights capturing generic debugging and optimization strategies. These insights are injected either as on-demand skills or as an upfront playbook. In our evaluation, the extracted insights improve the normalized private-test score by 123.9% over a no-insight baseline. Controlled experiments further show that the three insight types are complementary; the better injection mode (skill vs. playbook) varies by tasks, with skills more cost-effective overall; and the gains generalize across backend LLMs.