arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

限制非确定性:作为数据库管理系统的人工智能驱动研究系统,实现可靠、无浪费、透明和协作式研究 [愿景]

Confining Nondeterminism: AI-Driven Research Systems as DBMSs for Reliable, Non-Wasteful, Transparent, and Collaborative Research [Vision]

Kyoungmin Kim, Anastasia Ailamaki

arXiv 2607.10508首次发表:更新:

发表机构

EPFL(瑞士联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究指出大语言模型代理在研究中存在不可信问题,根源是代理循环的随机调用无法检查内部状态。借鉴数据库经验,建议用确定性带版本的数据流引擎组织研究项目,大语言模型作查询编译器,五条设计规则确保研究可靠、透明且协作。

AI 中文摘要

进行研究(提出想法、编写和运行代码、分析结果)的大语言模型代理已经能够将一项研究从研究问题推进到生成图表,但却无法得到完全信任。连续两次提出相同问题会得到不同答案;代理会宣布一个未通过任何执行产生的数字,且工具使用也无法阻止这种情况,因为没有任何东西将代理报告的内容与工具返回的内容绑定;上游的一个小变化会使下游结果悄然过时,且无法列出哪些结果过时;代理会重新运行预处理并重写它已经生成的代码。我们认为这些失败都有一个共同根源:当今代理循环中的每一步都是对大语言模型的随机调用,其内部状态无人能检查,包括代理自身。我们不是试图窥视大语言模型内部,而是从数据库中汲取经验,数据库无需被监视就能赢得信任,因为基于定义良好的状态的确定性操作使其保证得以构建。我们建议以同样的方式组织一个研究项目。该项目存在于一个确定性的、带版本的数据流引擎中(实际上是一个针对物化视图的查询计划),大语言模型与用户一起是一个随机编译器,只能编辑该计划。执行器从不调用大语言模型;大语言模型的输出仅作为带版本的代码和数据进入,执行器随后运行这些代码和数据,任何断言的结果只有在有执行支持的情况下才会进入记录。在这个边界处的五条设计规则将熟悉的数据库机制,从版本控制和出处到增量维护和基于成本的调度,转变为使研究可靠、无浪费、透明和协作的保证。本报告展示了诊断、要求和设计;保证演练、一个原型和研究议程将出现在准备中的完整版本中。我们认为,大语言模型应该是查询编译器,而绝不是执行器。

英文摘要

LLM agents that conduct research (proposing ideas, writing and running code, analyzing results) can already carry a study from research question to figures, yet cannot be fully trusted. The same question asked twice in a row returns different answers; the agent announces a number that no execution produced, and tool use does not prevent this, because nothing binds what the agent reports to what its tools returned; a small upstream change leaves downstream results silently stale, with no way to list which ones; and the agent re-runs preprocessing and rewrites code it has already produced. We argue these failures share one root: every step of today's agent loop is a stochastic LLM call whose internal state nobody, including the agent, can check. Rather than trying to see inside the LLM, we take a lesson from databases, which earn trust without being watched, because deterministic operators over well-defined state make their guarantees hold by construction. We propose organizing a research project the same way. The project lives in a deterministic, versioned dataflow engine (in effect, a query plan over materialized views), and the LLM, together with the user, is a stochastic compiler that may only edit that plan. The executor never calls the LLM; LLM output enters only as versioned code and data that the executor then runs, and any asserted result enters the record only with an execution behind it. Five design rules at this boundary turn familiar database machinery, from versioning and provenance to incremental maintenance and cost-based scheduling, into guarantees that make research reliable, non-wasteful, transparent, and collaborative. This report presents the diagnosis, the requirements, and the design; the guarantee walkthrough, a prototype, and the research agenda appear in the full version, in preparation. The LLM, we argue, should be the query compiler, never the executor.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑