WhatWorkedBench:AI智能体实验理解的基准测试
WhatWorkedBench: Benchmarking Experimental Understanding in AI Agents
- Carnegie Mellon University(卡内基梅隆大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出WhatWorkedBench基准,通过响应曲面预测和GP拟合提升AI智能体实验理解,覆盖36任务,显著提高效应恢复准确性。
AI中文摘要:
AI研究智能体需要可靠地了解其实验如何改变结果。我们引入WhatWorkedBench来衡量实验理解,即在预算有限的实验后对组件变化预测的准确性。智能体检查代码、选择测量,并提交一个响应曲面,即一个预测每种组件设置组合得分的表格。穷举CPU执行提供了在保持其他组件固定时改变每个组件的参考效应。这些效应涵盖了来自30个数据源和8种工作流类型的36个任务中的组合变化,包含1248条配置记录。核心评估结合了所有八个家族的4206条数值控制记录和原始六个家族的108个智能体片段。在八次新测量中,配对效应岭回归在22个源中的15个上选择了最优值,并在三个源上将每个效应误差限制在得分范围的10%以内。将高斯过程(GP)拟合到相同的智能体观测中,提高了效应恢复(相对于真实效应幅度的准确性),在原始Flash队列中从0.632提高到0.698,在另一个队列中从0.621提高到0.720。在六个完成的节拍检测和图形提交中,相同观测的GP将家族宏恢复从0.303提高到0.455。在六个具有六个二元选项的工作流中,在20次新测量下,编码代码等价性(即行为相同的配置)将GP恢复从0.248提高到0.462。WhatWorkedBench支持实验智能体、自适应实验设计、数值推断以及程序结构使用的研究。
英文摘要:
AI research agents must predict the effects of computational changes after budgeted experiments. WhatWorkedBench evaluates this experimental understanding through a delivered response surface of configuration scores. Agents inspect workflow code and buy measurements; exhaustive CPU references score conditional component effects, mean pair interactions, configuration choice, and delivery. The catalog spans 36 task conditions, 30 sources, 8 workflow families, and 1248 indexed configuration records, with 4,206 numerical controls. At eight purchased measurements plus two free anchors, pair-effect ridge selects an exact optimum on 15 of 22 sources; 13 of these cases have at least one conditional-effect error exceeding 10% of the task utility range. Shared-estimator comparisons measure acquisition and reconstruction on common observations. In a prospective typed study on 12 four-factor sources, DeepSeek Flash submits 12/12 direct tables and gains 0.147 recovery over a Gaussian process (GP) fitted to the same observations. Pro delivers 11/12 artifacts, with an all-attempt GP difference of -0.001 and a delivered-only difference of +0.063. On six prespecified new agent-evaluation sources, Flash and Pro gains are 0.149 and 0.074. In eight typed six-factor episodes, seven final tables satisfy verified code equivalences. Four fresh agents pass all six registered rules through named estimators and deliver consistent tables at 0.682 recovery versus 0.710 for separate direct-table runs. WhatWorkedBench links acquisition, inference, program structure, and delivery.