arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25770cs.MA

HypoForge:一种通过科学技能学习实现自动假设生成与测试的自改进多智能体框架

HypoForge: A Self-Improving Multi-Agent Framework for Automated Hypothesis Generation and Testing via Scientific Skill Learning

Ziqing Qian, Jiaying Lei, Yifang Wang, Nan Cao

首次发表
浏览论文内容

中文总结 AI 辅助

提出HypoForge多智能体框架,通过匹配阶段特定监督的技能学习策略,无需微调基础模型即可实现自改进,在相关基准上优于现有框架。

中文摘要 AI 辅助

大型语言模型(LLMs)已使AI科学家系统能够实现科学发现自动化,但现有方法大多依赖静态提示或固定工作流程,无法积累经验以实现持续改进。我们提出HypoForge,这是一种经验引导的多智能体框架,可学习可复用的科学技能以实现自动假设生成与测试。HypoForge基于以下观察:假设生成与测试两个阶段涉及不同的监督信号。对于无法获取明确反馈的假设生成,HypoForge采用对抗生成器-判别器机制,通过对比批判来改进推理;对于可获取实证反馈的假设测试,HypoForge从执行结果与真实结果中学习测试技能。通过使技能学习策略与阶段特定监督相匹配,HypoForge无需对基础模型进行微调即可实现持续改进。在假设生成与测试基准上的实验表明,HypoForge始终优于现有AI科学家框架及技能级变体,进一步分析也证明了所提出的阶段特定技能学习范式的有效性。

英文摘要

Large language models (LLMs) have enabled AI scientist systems to automate scientific discovery, yet existing approaches most rely on static prompting or fixed workflows and fail to accumulate experience for continual improvement. We propose HypoForge, an experience-guided multi-agent framework that learns reusable scientific skills for automated hypothesis generation and hypothesis testing. HypoForge is built on the observation that these two stages involve different supervision signals. For hypothesis generation, where explicit feedback is unavailable, HypoForge adopts an adversarial generator--discriminator mechanism to improve reasoning through comparative critique. For hypothesis testing, where empirical feedback is available, HypoForge learns testing skills from execution outcomes and ground-truth results. By matching skill learning strategies with stage-specific supervision, HypoForge enables continual improvement without fine-tuning foundation models. Experiments on hypothesis generation and testing benchmarks show that HypoForge consistently outperforms existing AI scientist frameworks and skill-level variants. Further analysis demonstrates the effectiveness of the proposed stage-specific skill learning paradigms.

↑