arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39392cs.AI

自主研究的实验经验建模

Experimental Experience Modeling for Autonomous Research

  • State Key Laboratory of AI Safety(人工智能安全国家重点实验室)
  • Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Baidu Inc.(百度公司)

机构由 AI 辅助整理,请以论文原文为准。

Wenda Wei, Yingchen Zhang, Ruqing Zhang, Jiafeng Guo, Daiting Shi, Xueqi Cheng

AI总结:

针对自主研究智能体实验成本高的问题,提出实验经验建模框架,通过获取、重用和积累实验经验来优化实验决策,在基准上提升性能并降低交互开销。

AI中文摘要:

自主研究智能体能够生成假设并进行实验,但实验仍然是计算成本的主要来源。一个基本挑战是决定哪些实验值得运行,尤其是在先前证据不足以消除不确定性的情况下。然而,当前的研究智能体在做出此类决策时缺乏系统性地利用实验经验的方法。我们引入了实验经验建模(EEM),这是一个通过获取、重用和积累实验经验来做出明智实验决策的框架。EEM从早期的实验轨迹中提取与决策相关的记录,将其提炼为可重用的经验,并将其组织在经验库中。对于新的实验决策,EEM检索相关的历史经验,并评估其是否为决定候选方向是否值得进一步投资提供了充分支持。当历史经验不足时,EEM会进行有针对性的低成本试点实验,按需获取缺失的决策相关经验。然后,它将新获取的经验与检索到的历史经验相结合,以确定该方向是否值得进行需要大量资源的全面评估。由此产生的实验结果被进一步提炼为可重用的经验,使经验库能够通过迭代积累不断增长。在自主研究基准上的实验表明,EEM提高了研究性能,同时减少了模型交互开销,证明了重用积累的经验以及仅在需要时获取额外经验的价值。

英文摘要:

Autonomous research agents can generate hypotheses and conduct experiments, but experimentation remains a major source of computational cost. A fundamental challenge is deciding which experiments are worth running, particularly when prior evidence is insufficient to resolve uncertainty. Yet current research agents lack a systematic way to leverage experimental experience when making such decisions. We introduce Experimental Experience Modeling (EEM), a framework for making informed experimental decisions by acquiring, reusing, and accumulating experimental experience. EEM extracts decision-relevant records from earlier experimental trajectories, distills them into reusable experience, and organizes them in an experience library. For a new experimental decision, EEM retrieves relevant historical experience and assesses whether it provides sufficient support for deciding whether a candidate direction warrants further investment. When historical experience is insufficient, EEM conducts a targeted, low-cost pilot experiment to acquire the missing decision-relevant experience on demand. It then combines this newly acquired experience with retrieved historical experience to determine whether the direction warrants full-scale evaluation, which requires substantial resources. The resulting experimental outcomes are further distilled into reusable experience, allowing the library to continually grow through iterative accumulation. Experiments on autonomous research benchmarks show that EEM improves research performance while reducing model interaction overhead, demonstrating the value of reusing accumulated experience and acquiring additional experience only when needed.

↑