AI 中文总结
针对机器学习开发智能体学习成本高且工具使用难以学习的问题,提出ToolMLBench工具套件和SPICE奖励方法,显著提升域内及域外任务成功率。
AI 中文摘要
机器学习工程(MLE)智能体已取得实质性进展,但通过机器学习实验进行学习在时间和计算上仍然代价高昂。合成环境降低了这些成本,同时引入了数据和实验设置上的变化,这些变化需要针对特定任务的诊断。仅提供诊断工具并不能确保智能体学会何时使用它们或如何根据其发现采取行动。我们引入了ToolMLBench,一套用于数据检查、代码验证和实验诊断的可执行工具,以及一个用于学习其使用的SFT和RL流水线。诊断调用获取的证据价值取决于后续决策,因此最终结果对应强化哪些调用提供的指导有限。我们通过SPICE解决了这一挑战,它衡量特权上下文如何改变采样工具动作的可能性,并将这种差异作为回合级奖励与最终结果一起使用。我们在80个合成任务上训练,并在25个域内和10个域外任务上评估。仅提供工具接口和描述在未适应的模型上产生不一致的收益。使用相同的诊断接口,我们的训练流水线将Qwen3-8B的域内成功率从24.8%提升到52.4%,将Qwen3.5-35B-A3B的域内成功率从35.6%提升到69.2%。后者在域外也从31%提升到48%,支持在保留的来源和目标上学习诊断工具使用。
英文摘要
Machine learning engineering (MLE) agents have made substantial progress, but learning through ML experimentation remains costly in time and computation. Synthetic environments reduce these costs while introducing variations in data and experimental settings that require task-specific diagnosis. Access to diagnostic tools alone does not ensure that agents learn when to use them or how to act on their findings. We introduce ToolMLBench, a suite of executable tools for data inspection, code verification, and experiment diagnosis, together with an SFT and RL pipeline for learning their use. Diagnostic calls acquire evidence whose value depends on subsequent decisions, so final outcomes provide limited guidance on which calls to reinforce. We address this challenge with SPICE, which measures how privileged context changes the likelihood of a sampled tool action and uses this difference as a turn-level reward alongside the final outcome. We train on 80 synthetic tasks and evaluate on 25 in-domain and 10 out-of-domain tasks. Providing tool interfaces and descriptions alone yields inconsistent gains across unadapted models. With the same diagnostic interface, our training pipeline raises in-domain success from 24.8% to 52.4% for Qwen3-8B and from 35.6% to 69.2% for Qwen3.5-35B-A3B. The latter also improves from 31% to 48% out-of-domain, supporting learned diagnostic tool use on held-out sources and targets.