发表机构
Zhejiang University; University of Electronic Science and Technology of China; Beijing University of Posts and Telecommunications; Zhejiang University of Technology(浙江大学; 电子科技大学; 北京邮电大学; 浙江工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出AutoSciRub框架,在自主科学研究前生成任务特定可执行评分准则,经多基准测试可提升智能体研究表现,为自主科研提供有效控制机制。
AI 中文摘要
自主科学研究智能体正越来越多地应用于从文献综述、数据分析、实验到报告生成的端到端科学工作流程中。然而,开放式研究任务往往未明确规定完成任务所需的分析、方法和成功标准,导致智能体可能遗漏重要分析、使用不当方法或得出证据不足的结论。为解决该问题,本文提出AutoSciRub,这是一种评估优先的框架,会在研究执行前生成任务特定的可执行评分准则,并利用该准则指导执行、准则级验证及迭代修订。AutoSciRub将未明确的指令分解为原子科学目标,将其与相关文献和任务可见数据关联,合成具体、可操作且可验证的准则。生成的评分准则将隐含的实验和证据要求明确化,为实验和分析提供指导。在修订阶段,评分准则引导的验证会识别未满足的准则,从而实现对研究报告及其支撑产物的针对性优化。在ResearchClawBench上,AutoSciRub在固定Codex工具链下的三个骨干大语言模型(LLM)配置中均实现提升,平均增益为2.08分;在使用固定DeepSeek-V4-Flash骨干的三个智能体工具链中,平均增益为2.95分。在AstaBench端到端发现任务随机抽取的20个任务子集上,AutoSciRub在三个智能体工具链中平均提升16.8分,同时保持或增加成功完成的任务数量。这些结果表明,评估优先的指导为自主科学研究提供了有效且可泛化的控制机制(代码:this https URL)。
英文摘要
Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-ended research tasks often do not clearly specify the analyses, methods, and success criteria required to complete the task. As a result, agents may miss important analyses, use inappropriate methods, or draw conclusions that are insufficiently supported by evidence. To address the problem, we present AutoSciRub, an evaluation-first framework that induces a task-specific executable rubric before research execution, and uses it to guide execution, criterion-level verification as well as iterative revision. AutoSciRub decomposes an underspecified instruction into atomic scientific goals, grounds them in relevant literature and task-visible data, and synthesizes specific, actionable, and verifiable criteria. The resulting rubric makes implicit experimental and evidential requirements explicit, providing guidance for experiments and analyses. During revision, rubric-guided verification identifies unmet criteria and enables targeted refinement of the research report and its supporting artifacts. On ResearchClawBench, AutoSciRub consistently improves all tested configurations, with an average gain of 2.08 points across three backbone LLMs under the fixed Codex harness and 2.95 points across three agent harnesses using a fixed DeepSeek-V4-Flash backbone. On a randomly sampled 20-task subset of AstaBench E2E Discovery, AutoSciRub further achieves an average improvement of 16.8 points across three agent harnesses, while maintaining or increasing the number of successfully completed tasks. These results demonstrate that evaluation-first guidance provides an effective and generalizable control mechanism for autonomous scientific research (Code: https://github.com/zjunlp/AutoSciRub).
CommentsWork in progress