AI 中文总结
针对开放式研究任务,提出OEB基准,仅从智能体执行记录构建认知事件图,从证据、实验、修正等维度评分,发现119次运行中仅16-29%的声称改进为真。
AI 中文摘要
智能体越来越多地被赋予开放式研究任务:从自行设计的实验中总结出经验定律、改进一个无人知晓最优解的启发式算法,或打破一项历史纪录。它们的执行日志记录了研究的每一步,但这些运行仍仅通过结果分数来评判。仅凭该分数无法证明智能体的主张是否源于已执行的实验,且参考答案可能并不存在。我们评估智能体的认知过程:它如何形成假设、检验假设,并根据证据修正假设。我们提出OEB(开放结局基准),一种与基准无关的方法,它只读取智能体的执行记录,从不参考答案或结果分数。OEB将记录整合为一个统一的认知事件图,其边连接智能体陈述的命题与检验这些命题所执行的动作;每个节点携带一个精确的摘录,由代码对照记录进行验证。评分遵循一条原则:文字可以陈述命题,但只有执行动作返回的证据才能支持或反驳它,因此OEB检查智能体所写内容与其实际运行内容是否一致。基于该图,OEB从四个能力维度(证据、实验、修正、无奖励黑客行为)评分,主要衡量智能体在合理研究机会中的占比,并刻画六种描述智能体研究习惯的主观人格特质。我们对来自三个基准(LLM后训练、芯片设计、训练速度纪录)的12个任务中的119个现有运行进行评分。与已记录结果相比,智能体声称的改进中仅有16%至29%是真实的。在10个任务中的9个中,最佳运行在后半段尝试的新想法多于最差运行。人格特质读数与模型相关:对于每种特质,运行该模型的解释方差占比高于任务(中位数为43%对7%)。
英文摘要
Agents are increasingly given open-ended research tasks: discovering an empirical law from self-designed experiments, improving a heuristic whose optimum nobody knows, or beating a standing record. Their execution logs record every step of this research, yet the runs are still judged by their outcome score. That score alone does not establish whether an agent's claims follow from executed experiments, and a reference answer may be unavailable. We evaluate the agent's epistemic process: how it forms hypotheses, tests them, and revises them in response to evidence. We introduce OEB (Open-Endedness Bench), a benchmark-agnostic methodology that reads only the agent's execution record and never a reference answer or an outcome score. OEB compiles the record into a unified epistemic event graph whose edges connect the propositions the agent states to the executed actions that test them; each node carries an exact excerpt that code verifies against the record. One principle governs scoring: prose can state a proposition, but only evidence returned by an executed action can support or refute it, so OEB checks what the agent writes against what it actually ran. From the graph, OEB scores four competence axes (evidence, experiment, revision, and no reward hacking), mostly as the share of opportunities for sound research that the agent took, and profiles six subjective persona traits that describe the agent's research habits. We score 119 existing runs over 12 tasks from three benchmarks: LLM post-training, chip design, and a training-speed record. Against logged results, only 16-29% of the improvements agents claim are real. On 9 of 10 tasks, the best run tries more new ideas in its second half than the worst run. The persona readings follow the model: for every trait, the model that ran explains more of its variance across runs than the task (a median of 43% against 7%).
Comments18 pages, 7 figures. Code: https://github.com/ARA-Labs/oeb . Data: https://huggingface.co/datasets/AgentNativeResearchLab/oeb-scored-runs