arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.07611cs.AIcs.CL

AgentIdeaBench:智能体时代科学构思的基准测试

AgentIdeaBench: Benchmarking Scientific Ideation in the Agent Era

Yunxiang Mo, Tianshi Zheng, Yisen Gao, Rui Wang, Newt Nguyen Kim Hue Nam, Kelvin Kiu Wai Tam, Jiaxin Bai, Yangqiu Song, Ginny Wong, Simon See

首次发表
浏览论文内容

中文总结 AI 辅助

AgentIdeaBench提出多学科基准,在静态观察与主动探索两种设置下评估33个LLM的科学构思能力,发现主动探索带来约两倍性能提升且受能力门槛限制,为智能体时代科学构思提供测量基础。

中文摘要 AI 辅助

科学构思是从科学证据中提出新颖且可检验假设的能力,自主人工智能科学家依赖这一能力。现有评估大多通过要求模型从一组静态、精选的参考论文中生成想法来进行评估。这种被动设置与现代人工智能科学家的检索与推理工作流程相脱节,并且随着模型性能的提升,其区分度逐渐降低。我们引入了AgentIdeaBench,这是一个多学科基准,在两种匹配的设置下评估科学构思能力:静态观察和主动探索。我们报告了33个大型语言模型在跨越五个学科的40个密集评分子领域中的匹配静态-主动评估结果,采用了一个多维、经文献验证的评分框架,其评判者根据检索到的先前成果评估原创性。主动探索揭示了显著更大的能力提升空间,且该空间在模型间分布不均。性能提升速度约为静态观察下的两倍,且探索增益受能力门槛限制,更有利于最强模型而非最弱模型。该增益反映了更好的基础依据,提高了可行性、清晰度和特异性,而在我们的评判者看来,测得的原创性保持不变。我们进一步探索了科学世界建模,这是一种生成时循环,通过结构化思想实验来完善草拟假设。它对中等能力模型有益,而其影响在那些似乎已内化此类推理模式的前沿模型中有所减弱。AgentIdeaBench为未来科学构思研究提供了适合智能体时代的测量基础。

英文摘要

Scientific ideation is the capacity to formulate novel and testable hypotheses from scientific evidence, and autonomous AI scientists depend on it. Existing evaluations largely assess it by asking models to generate ideas from a static, curated set of reference papers. That passive setup departs from the retrieval-and-reasoning workflow of modern AI scientists, and it becomes less discriminative as models improve. We introduce AgentIdeaBench, a multidisciplinary benchmark that evaluates scientific ideation under two matched settings, static observation and active exploration. We report matched Static-Active evaluations for 33 LLMs across 40 densely scored subfields spanning five disciplines, using a multidimensional, literature-verified scoring framework whose critics assess originality against retrieved prior art. Active exploration reveals considerably more capability headroom, and that headroom is unevenly distributed across models. Performance scales about twice as fast as under static observation, and the exploration gain is capability-gated, favoring the strongest models over the weakest. The gain reflects better grounding, improving feasibility, clarity, and specificity while leaving measured originality unchanged under our critics. We further explore Scientific World Modeling, a generation-time loop that refines a draft hypothesis through structured thought experiments. It benefits mid-capability models, and its impact diminishes among frontier models that appear to have internalized such reasoning patterns already. AgentIdeaBench gives future work on scientific ideation a measurement basis suited to the agent era.

发表机构

  • NVIDIA(英伟达)
  • NVIDIA AI Technology Center (NVAITC)(英伟达人工智能技术中心)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑