审计发现主张:面向智能体科学的双边准则,其负面侧可判定
Auditing Discovery Claims: A Two-Sided Criterion for Agentic Science, with the Negative Side Decidable
- University of Macau(澳门大学)
- South China University of Technology(华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究提出面向智能体科学的双边审计准则,通过实验揭示AI科学系统能力主张的夸大问题,证实智能体编写的程序在低计算量下可击败人类程序,但未明确可迁移的机制。
AI中文摘要:
当一个自我改进的AI科学系统宣称具备新能力时,证据通常是基准差值、描述长度阈值或p值,但这些都无法区分真实增益、额外搜索、验证器变更或对不可靠预言机的适应。我们构建了一个双边审计,其负面侧是一个形式化事实:无假结的预言机可证明无法表示交叉碱基对,因此先验验证器的范围在任何运行前就已确切、离线地被界定。“新”是相对于智能体的先前自我而言,而非相对于基础模型。首先,研究单个不可靠预言机能将能力主张夸大多少:我们发明了一种无求解器的算子,在其优化的预测器下解决了60个交叉RNA靶标中的43个,远超0/60的上下文无关基线;在三种预测器下,仅1/60留存。在相同的43个靶标上配对测试,该算子从未见过的预测器证实其2个设计优于最小自由能求解器的26个(p=8e-7),而系统及其自身预言机计算的任何统计量都未发现这一差距。其次,在无客观目标偏袒的评判下,智能体编写的程序可击败人类编写的程序,且计算量仅为后者的一小部分:在六个前沿模型中,两个算子未超时运行的模型,在951个配对单元(按靶标聚类,差值为[+0.108, +0.297],p=5e-5)上的迁移性能为0.293,而我们的模型为0.095,同时预言机调用量减少了4.6至10倍。研究还设定了三个层级:外部裁决者下的差异(已达成)、未被计算量收买的差异(双向均达成)、机制明确且可迁移(未达成,测试了七个候选,均未改变统计量)。上限是裁决面板本身:其三个预测器共享最近邻热力学参数,两个在κ=0.673处一致。审计对我们自身的系统同样严苛:匹配的无向搜索结果为0,无搜索探针显示84%的核心效应来自随机序列已能解决的靶标。
英文摘要:
When a self-improving AI-for-science system claims a new capability, the evidence is usually a benchmark delta, a description-length gate, or a p-value. None separates a real gain from extra search, from a changed verifier, or from adaptation to a fallible oracle. We build a two-sided audit whose negative side is a formal fact: a pseudoknot-free oracle provably cannot represent a crossing base pair, so the prior verifier's range is bounded exactly, offline, before any run. "New" is relative to the agent's prior self, never to the base model. First, how far a single fallible oracle can inflate a capability claim. An invented, solver-free operator solves 43/60 crossing RNA targets under the predictor it optimizes, above a context-free floor of 0/60; under three predictors, 1/60 survives. Paired on the same 43 targets, a predictor the operator never saw confirms 2 of its designs against 26 for a minimum-free-energy solver (p = 8e-7). No statistic computed from the system and its own oracle sees that gap. Second, agent-written procedures can beat a human-written one under a judge no objective can flatter, at a fraction of the compute. Of six frontier models, the two whose operators ran without timeouts carry over at 0.293 against our 0.095 (n = 951 paired units, target-clustered [+0.108, +0.297], p = 5e-5) while spending 4.6-10x fewer oracle calls. Three rungs: difference under an outside adjudicator (reached), not bought with compute (reached, both directions), mechanism identified and transferable (not reached; seven candidates tested, none moves the statistic). The ceiling is the panel itself: its three predictors share nearest-neighbour thermodynamic parameters, two agreeing at kappa = 0.673. The audit is as unsparing about our own system: matched undirected search is an exact zero, and a search-free probe puts 84% of our headline effect on targets a random sequence already solves.