AutoResearch:输入洞见,输出无幻觉
AutoResearch: Insight In, Hallucination Out
- EvoMap(伊沃地图)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
AutoResearch是两阶段自主研究系统,通过创意生成与执行结合,在多场景提升研究进展,修正实验结果,在RSICD基准上召回率提升且问题事件更少。
AI中文摘要:
自主研究系统越来越能够执行长周期研究工作流,但仅靠自动化无法确保研究过程保持科学严谨性。我们提出AutoResearch,这是一个连接创意生成与创意执行的两阶段系统,用于解决研究创意如何形成以及如何通过实验可靠验证的问题。在创意生成阶段,AutoResearch持续整合新兴研究信号与积累的领域知识,识别可迁移的机制性洞见,并利用多模型生成与交叉评审生成严谨、可测试的研究计划。在创意执行阶段,协调智能体将这些计划分解为实验,迭代执行并诊断实验结果,在接受研究结论前采用独立的循证评审。在跨模态检索、系统优化和基准驱动的机器学习等代表性场景中,AutoResearch将生成的创意转化为可衡量的进展,检测并修正不可靠的实验结果,并做出基于证据的决策以继续、修订或终止研究方向。例如,在RSICD基准上,AutoResearch生成的创意将平均召回率从32.84提升至34.69,同时仅记录5次经评审确认的问题事件,而其他自主研究系统的该数值为11至27次。这些结果表明,该研究过程在实验前已确立有意义的洞见,在接受结论前已确立结论的严谨性:输入洞见,输出无幻觉。
英文摘要:
Autonomous research systems are increasingly capable of executing long research workflows, yet automation alone does not ensure that the resulting process remains scientifically grounded. We introduce AutoResearch, a two-stage system that connects Idea Generation with Idea Execution to address both how research ideas are formed and how they are reliably established through experimentation. In Idea Generation, AutoResearch continuously integrates emerging research signals with accumulated domain knowledge, identifies transferable mechanistic insights, and uses multi-model generation and cross-review to produce grounded, testable research plans. In Idea Execution, coordinated agents decompose these plans into experiments, iteratively implement and diagnose them, and employ independent evidence-based review before accepting research conclusions. Across representative settings in cross-modal retrieval, systems optimization, and benchmark-driven machine learning, AutoResearch turns generated ideas into measurable progress, detects and corrects unreliable experimental results, and makes evidence-conditioned decisions to continue, revise, or terminate research directions. For example, on RSICD benchmark, an AutoResearch-generated idea improves mean Recall from 32.84 to 34.69, while recording only 5 audit-confirmed issue events compared with 11-27 for other autonomous research systems. These results demonstrate a research process in which meaningful insight is grounded before experimentation and conclusions are grounded before acceptance: Insight In, Hallucination Out.