arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ULTRADISCOVERY:在相互关联、认识论开放宇宙中的溯因探索

ULTRADISCOVERY: Abductive Exploration in an Interconnected, Epistemically Open Universe

Weihan Li, Tianshi Zheng, Yangqiu Song, Ginny Y. Wong, Simon See

arXiv 2610.03092首次发表:更新:

发表机构

The University of Tokyo; HKUST; NVIDIA(东京大学; 香港科技大学; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ULTRADISCOVERY通过2×2设计分离表征开放与证据分散需求,发现现有模型在开放表征下无法引入新实体或重写变量,且难以做出精确跨域预测,揭示溯因探索的核心困难在于将证据组合为可迁移表征。

AI 中文摘要

科学发现往往始于零散线索呼唤一种描述世界的新方式。当世界在认识论上是开放的时,这种溯因探索可能要求构建解释所依托的表征;当世界在结构上相互关联时,则需整合分散在不同情境中的证据。现有基准很少将这两种需求分开或独立控制。我们引入ULTRADISCOVERY,一个包含五个领域的交互式世界,其中智能体修正一个最初有效的理论,并预测一个未见过的跨领域干预的结果。一个2×2设计使得表征保持开放或予以披露,并使得证据保持分散或予以对齐,而潜在动力学固定不变。在表征开放的情况下,十一个模型中的智能体经常撤回它们被教导的公理,且没有一个引入未观察到的实体或重写替换所需变量。披露使干预请求增加两倍,并增加了世界所提供的十八个发现中的约一个,而对齐增加较少。两个供应商框架系统将发现带入更多领域,其中一个在开放情节中重写变量。没有系统在200次付费行动内做出精确预测。在更大预算下,一个精确预测在两种辅助条件下出现,而每个开放情节仍不精确。结果将困难定位于从积累证据到将其组合成可迁移表征的步骤。

英文摘要

Scientific discovery often begins when scattered clues call for a new way of describing the world. Such abductive exploration can require constructing the representation in which an explanation is stated, when the world is epistemically open, and composing evidence scattered across contexts, when it is structurally interconnected. Existing benchmarks rarely separate these two demands or control them independently. We introduce ULTRADISCOVERY, an interactive world of five domains in which an agent revises an initially successful theory and predicts the outcome of an unseen cross-domain intervention. A $2 \times 2$ design leaves the representation open or discloses it, and leaves the evidence distributed or aligns it, with the latent dynamics fixed. With the representation open, agents across eleven models often retract the axiom they were taught, and none introduces the unobserved entity or rewrites the variables that a replacement requires. Disclosure triples intervention requests and adds about one of the eighteen findings the world affords, and alignment adds less. Two vendor-harness systems carry discovery into more domains, and one of them rewrites the variables in Open episodes. No system makes the exact prediction within 200 paid actions. At larger budgets one exact prediction appears with both aids, while every Open episode remains inexact. The results locate the difficulty in the step from accumulating evidence to composing it into a representation that transfers.

Comments47 pages, 19 figures, 15 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑