发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TraceML通过构建人类与智能体在Kaggle竞赛中的配对轨迹数据集,实证分析了人类与智能体在机器学习开发规划上的差异,发现智能体存在行为循环问题,提炼的人类规划提示可缩小部分差距。
AI 中文摘要
大型语言模型能为孤立问题写出正确代码,但在自主机器学习开发方面仍远弱于人类,自主机器学习开发中智能体需在数小时的反馈中修改数据管道、模型和验证流程,且在大多数竞赛中仍落后于优秀人类参赛者。基于结果的基准记录了这一差距,但未记录其原因,因为它们仅对最终提交结果进行评分,丢弃了背后的开发过程。我们推出TraceML,它在同一版本级架构下配对人类和智能体在竞赛中的工作:涵盖134场竞赛的4465条人类Kaggle轨迹,其中7场竞赛也有两个智能体框架参与,共得到430条配对人类轨迹和207条智能体轨迹。每个代码版本都带有其分数、时间戳,以及所采取的操作、意图、编辑规模和分数影响的标签。从这个角度看,差距变得具体:专家交替进行数据处理、验证、模型修改和集成,并会重新采用曾搁置的方法;而每个智能体框架则陷入狭窄循环:Codex将步骤用于重新加权集成和调整提交,MLEvolve原地变异其模型,两者均未以人类的速率转变方向,也未重新开启被放弃的工作。从人类实践中提炼的简短规划提示,能使其命名的行为向人类特征靠拢并提升分数,但努力特征仍呈智能体形态:指令仅缩小了可归因于指令的那部分差距。我们在https URL发布了该语料库、架构、标注工具和提取管道。
英文摘要
Auto-research agents now run machine-learning development unattended for hours, revising data pipelines, models, and validation from their own feedback, yet on most competitions they still finish below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We introduce TraceML, which pairs human and agent work on the same competitions under one version-level schema: 4{,}465 human Kaggle trajectories across 134 competitions, seven of which are also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its intent, the edit size, and the score effect. Read this way, the gap becomes concrete. Experts alternate data work, validation, model changes, and ensembling, and return to approaches they had set aside. Each agent scaffold instead collapses into a narrow loop: Codex spends its steps re-weighting ensembles and tuning submissions, MLEvolve mutates its model in place, and neither pivots at the human rate nor reopens abandoned work. A short planning prompt distilled from human practice moves the behaviors it names toward the human profile and lifts scores, but the effort profile stays agent-shaped: instruction closes only the part of the gap that reduces to instructions. We release the corpus, the schema, and the labeling models, together with an open-source toolkit that turns a run from any command-line agent into a TraceML trajectory and reads it against the human cohorts.