arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AgentDiscover:以最小搜索脚手架实现自主发现

AgentDiscover: Autonomous Discovery with Minimal Search Scaffolding

Mahdi Farahbakhsh, Ilan Sela, Fatemeh Doudi, Vishnu Teja Kunde, Krishna Narayanan, Jean-Francois Chamberland, Dileep Kalathil

arXiv 2610.05334首次发表:更新:

发表机构

Texas A&M University(得克萨斯农工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

AgentDiscover通过让编码代理自主规划搜索并利用数据库作为长期记忆,以最小脚手架实现高效科学发现,在多项任务上超越现有框架并降低成本。

AI 中文摘要

用于科学发现的大语言模型框架通常依赖于一个固定的、由人类设计的算法,该算法决定模型在每一步看到什么,模型仅扮演提议者的角色。模型除了被展示的内容外,对搜索一无所知。随着模型能力增强,一个问题浮现:由人类在运行前选择的搜索策略,是否比将模型从提议者提升为规划者并让其掌控搜索更具可扩展性?苦涩教训表明,预先选择策略正是那种通用方法最终会超越的手工设计结构。我们提出AgentDiscover,其中编码代理利用其上下文作为工作记忆来规划搜索、运行实验,并将每次尝试记录在包含想法、候选及其关系的数据库中。该数据库作为代理的长期记忆,其结构使得经典算法如MAP-Elites和蒙特卡洛树搜索的选择规则均简化为单个查询,代理可自由使用、组合或替换这些查询。服务器维护数据库并在每次提交后引导代理,使其在长时间运行中保持方向。在实验中,AgentDiscover比现有框架更具成本效益,以更低成本达到更优分数。在核工程、生物学、算法设计和数学任务上,AgentDiscover优于先前的发现框架。其程序在七场过去的AtCoder启发式竞赛中本可位列人类参赛者之首,在十一个数学和系统优化任务上,它匹配或超越了所有使用相同模型的基线。我们的代码可在该https URL获取。

英文摘要

Frameworks that use large language models for scientific discovery typically rely on a fixed, human-designed algorithm that decides what the model sees at each step, leaving the model only the role of proposer. The model knows nothing of the search beyond what it is shown. As models grow more capable, a question arises: does a search strategy chosen by a human before the run scale better than promoting the model from proposer to planner and letting it own the search? The Bitter Lesson suggests that choosing the strategy in advance is the kind of hand-designed structure that general methods eventually outscale. We introduce AgentDiscover, in which a coding agent plans the search using its context as working memory, runs experiments, and records every attempt in a database of ideas, candidates, and their relations. This database serves as the agent's long-term memory and is structured so that the selection rules of classical algorithms such as MAP-Elites and Monte Carlo tree search each reduce to a single query, which the agent is free to use, combine, or replace. A server maintains the database and steers the agent after every submission, keeping it on course over long runs. In our experiments, AgentDiscover is more cost-efficient than existing frameworks, reaching better scores at lower cost. On tasks in kernel engineering, biology, algorithm design, and mathematics, AgentDiscover outperforms prior discovery frameworks. Its programs would have placed first among human competitors in seven past AtCoder heuristic contests, and on eleven mathematical and systems optimization tasks it matches or exceeds every baseline that uses the same model. Our code is available at https://github.com/mhdfb/AgentDiscover.

Comments23 pages, 8 figures, 11 tables. Code: https://github.com/mhdfb/AgentDiscover

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑