arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25421cs.PLcs.AI

超越自然语言:面向自主科学的智能体原生语言

Beyond Natural Language: An Agent-Native Language for Autonomous Science

  • ARA Lab(ARA实验室)

机构由 AI 辅助整理,请以论文原文为准。

Yifeng He, Jiachen Liu

AI总结:

本文提出Lara,一种可机器检查的语言,将研究论证转化为可执行工件,实现自主科学中主张的自动化验证与审计,支持毫秒级状态重检,并在Lean 4中机械化元理论。

AI中文摘要:

随着自主AI智能体承担科学探究的每个阶段,研究产出的规模已远超人类审查能力。然而,科学交流仍依赖自然语言散文:这种非正式媒介容易产生歧义、隐藏假设和无法追踪的局限性,机器无法可靠地审计。我们提出Lara,一种可机器检查的语言和协议,用于检查和修订研究主张的支持。通过将研究论证转化为可执行工件,Lara为自主科学提供了认知内核:它为研究智能体启用自动化验证流水线,让声明的桥接将跨论文的论证连接成可审计的网络,并允许人类和机器在毫秒内重新检查编码主张的状态。在Lara程序中,作者明确声明其主张、支持证据和假设,以及已知的反对意见或局限性。一个轻量级、确定性的检查器裁决这些交互,为每个主张分配可复现的状态:“已证明”、“被击败”、“有争议”或“缺口”,后者标记支持不完整的主张并定位未解答的问题。案例研究涵盖实证审查、无测量的哲学辩论,以及当假设的公理被撤销时支持的丧失。我们建立了主张检查和跨上下文论证传输的元理论,并在Lean 4(约117,000行)中机械化语义保证,留下三个论证在纸上。经审计的公共元理论无“sorry”且仅使用Lean的三个标准公理;一些可执行示例额外信任原生求值。

英文摘要:

As autonomous AI agents take on every stage of scientific inquiry, research output is expanding far beyond human review capacity. Yet scientific communication still relies on natural-language prose: an informal medium prone to ambiguity, hidden assumptions, and untracked limitations that machines cannot reliably audit. We introduce Lara, a machine-checkable language and protocol for checking and revising support for research claims. By turning research arguments into executable artifacts, Lara provides an epistemic kernel for autonomous science: it enables automated validation pipelines for research agents, lets declared bridges connect arguments across papers into an auditable network, and allows both humans and machines to recheck the standing of an encoded claim in milliseconds. In a Lara program, authors explicitly declare their claims, supporting evidence and assumptions, and known objections or limitations. A lightweight, deterministic checker adjudicates these interactions, assigning each claim a reproducible status: "justified", "defeated", "contested", or "gap", which marks a claim whose support is incomplete and locates the unanswered question. Case studies cover empirical review, a philosophical debate without measurements, and the loss of support when an assumed axiom is withdrawn. We establish the metatheory of claim checking and cross-context argument transport, and mechanize the semantic guarantees in Lean 4 (roughly 117,000 lines), leaving three arguments on paper. The audited public metatheory is "sorry"-free and uses only Lean's three standard axioms; some executable examples additionally trust native evaluation.

↑