arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向可审计的人工智能科学家:大语言模型智能体的假设演化协议

Toward Auditable AI Scientists: A Hypothesis Evolution Protocol for LLM Agents

Izumi Takahara, Teruyasu Mizoguchi

arXiv 2607.09195首次发表:更新:

发表机构

Institute of Industrial Science, The University of Tokyo(东京大学产业科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究旨在让大语言模型智能体在科学发现中发挥核心作用。提出假设演化协议(HEP),使假设生成等操作可审计。在材料科学任务中,该协议能让智能体运行关键循环,跨问题概括并随模型能力提升更充分利用协议,迈向可审计的人工智能科学家。

AI 中文摘要

大语言模型智能体在人工智能驱动的科学发现中有望发挥核心作用,它们具备广泛知识、灵活推理和工具使用能力,有潜力通过反复提出假设、测试并根据证据修正信念来自主探索和解决科学问题。然而当前智能体中,这些假设、测试和信念更新都隐藏在无结构日志中,缺乏审计机制。本文提出假设演化协议(HEP),作为一种智能体框架,提供明确、可审计的假设生成、评估和演化操作。在材料科学研究任务中,配备HEP的智能体运行规划式智能体所缺乏的假设-测试-证据-信念循环,能跨研究问题进行概括,并随着基础大语言模型能力增强更充分地利用该协议。这些结果朝着可审计的人工智能科学家迈进了一步,其科学推理可被检查、验证和拓展。

英文摘要

Large language model (LLM) agents are increasingly expected to play a central role in AI-driven scientific discovery. Equipped with broad knowledge, flexible reasoning, and tool use, they have the potential to autonomously explore and solve scientific problems by repeatedly proposing hypotheses, testing them, and revising their beliefs in the light of the evidence. In current agents, however, these hypotheses, tests, and belief updates are buried in unstructured logs, and no mechanism lets the agent or the human researcher audit that process. Here we propose the Hypothesis Evolution Protocol (HEP), an agent harness that provides hypothesis generation, evaluation, and evolution as explicit, auditable operations. On materials-science research tasks, a HEP-equipped agent operates the hypothesis--test--evidence--belief cycle that planning-style agents lack, generalizes across research questions, and exploits the protocol more fully as the base LLM becomes more capable. These results mark a step toward auditable AI scientists, whose scientific reasoning can be inspected, verified, and built upon.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑