arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22375cs.AI

IDEAgent:用于研究想法生成的智能体质量多样性搜索

IDEAgent: Agentic Quality-Diversity Search for Research Idea Generation

Varun Gumma, Navonil Majumder, Soumitra Sinhahajari, Soujanya Poria

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对大语言模型在生成研究想法时质量与多样性独立导致的局限,提出将研究构思视为质量多样性搜索。介绍多智能体框架IDEAgent,通过多目标反馈驱动质量,轻量级记忆等实现多样性,开发Yield指标评估,实验表明其性能优于基线,还证实修复改进对质量提升的重要性并开源。

中文摘要 AI 辅助

在过去几年中,大语言模型显著地使科学发现过程自动化。然而,现有系统存在一个核心局限:它们在质量或多样性方面独立地生成和优化想法,这常导致生成的想法彼此接近,或产生大量琐碎、不合理或不清晰的概念。在这项工作中,我们认为研究构思应被视为两个目标的结合,并构建为质量多样性(QD)搜索。为此,我们引入了IDEAgent,这是一个通过谱系管理想法演化的多智能体框架。我们使用多目标反馈联合驱动质量以进行专门的修复和改进,而多样性则通过轻量级顺序记忆以及与已完成想法、其历史祖先和被拒绝的提议进行显式比较来实现。为了系统地评估这种QD结合,我们开发了Yield,这是一个联合指标,用于计算满足预定质量阈值的最大一组相互不同的想法。最后,通过对计算机科学8个领域的32个主题的评估,我们表明IDEAgent在Yield上比最佳基线性能高出3.89倍,同时在更多主题上实现了非零Yield。我们还通过质量改进分析进一步证实了这些发现,表明修复和改进对于建立逻辑严谨性和清晰度同时保持非显而易见性至关重要。为鼓励未来基于QD搜索的构思研究,我们在这个https URL上开源了IDEAgent。

英文摘要

Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize ideas independently for either Quality or Diversity. This often leads to the generation of ideas in close proximity to one another or to a large set of trivial, unsound, or unclear concepts. In this work, we instead argue that research ideation should be treated as a conjunction of both objectives and framed as a Quality-Diversity (QD) search. In line with this perspective, we introduce IDEAgent, a multi-agent framework that manages the evolution of ideas through lineages. We jointly drive Quality using multi-objective feedback for dedicated repair and refinement, while Diversity is achieved through lightweight sequential memory and explicit comparison against completed ideas, their historical ancestors, and rejected proposals. To systematically evaluate this QD conjunction, we develop Yield, a joint metric that computes the largest set of mutually diverse ideas that satisfy a predetermined quality threshold. Finally, through evaluations across 32 topics spanning 8 domains of Computer Science, we show that IDEAgent outperforms the best baseline by 3.89x on Yield, while achieving non-zero Yield on 8x more topics. We further corroborate these findings through an analysis of quality improvements, showing that repair and refinement are crucial for building logical rigor and clarity while preserving non-obviousness. To encourage future research on QD-search-based ideation, we open-source IDEAgent at https://github.com/declare-lab/IDEAgent.

发表机构

  • DeCLaRe Lab, Nanyang Technological University(声明实验室,南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑