arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

模型何时可替代实验?代理驱动设计中的审计、许可与信任代价

When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design

Shuangxiu, Ma, Wenhe, Zhao

arXiv 2608.01378首次发表:更新:

发表机构

The Ohio State University(俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究明确了机器学习代理替代实验的安全条件,推导了模型作为神谕的最小准则,提出选择感知审计可降低认证评估成本,经审计的筛选使认证神谕成本降低25倍。

AI 中文摘要

化学、材料科学与机器学习领域的设计活动存在共同瓶颈:确定候选的真实性能需开展昂贵评估——包括实验、第一性原理模拟或完整训练过程。机器学习代理(surrogate)可预测上述结果,如今不仅用于提出候选,还用于对其评分,甚至将自身预测当作测量值反馈到搜索过程中。通过数学分析及三项经详尽基准验证的设计任务验证,本文明确了该实践的安全条件、任何安全证书的必然代价,以及该替代方案何时具备可证优势。预测精度无法作为信任依据:接近完美的R²可能伴随最差候选选择,且筛选N个候选会使选中候选的过度预测产生可量化的“选择税”,存在匹配的上下界。安全条件源于架构规则:预测可不受限制地用于提出和训练,但所有经认证的结论必须基于真实评估;该规则对代理无假设要求,且是必要条件,因为若允许预测以测量值的身份参与认证,会引发确定性自确认失效模式。本文推导了模型可作为“神谕(oracle)”的最小准则(秩保持而非精度),证明信任需通过选择感知审计获得,该审计在查询复杂度上最优,且证明了经审计的代理何时可降低经认证评估成本的二分性。在6种任务场景下的432次代理拟合实验中,审计统计量与部署搜索性能的斯皮尔曼秩相关系数为0.80-0.99,而R²与部署遗憾的秩相关系数低至0.33;经审计的筛选将经认证神谕成本降低了实测的25倍。

英文摘要

Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run. Machine-learning surrogates that predict these outcomes are increasingly used not only to propose candidates but to grade them, and even to feed their own predictions back into the search as though they were measurements. Through mathematical analysis validated on three exhaustively ground-truthed design tasks, we establish when this practice is safe, what any certificate of safety must cost, and when the substitution provably pays. Predictive accuracy cannot anchor trust: near-perfect R^2 is compatible with worst-possible selections, and screening N candidates inflates the over-prediction at the selected candidate by a quantifiable "selection tax" with matching upper and lower bounds. Safety follows instead from an architectural rule - predictions may propose and train without restriction, but every certified conclusion must rest on true evaluations - which is sufficient with no assumptions on the surrogate, and necessary, since admitting predictions into certification with the standing of measurements opens a deterministic self-confirmation failure mode. We derive the minimal criterion under which a model may act as an oracle (rank preservation, not accuracy), show that trust must be purchased through selection-aware audits that are optimal in query complexity, and prove a dichotomy fixing when audited surrogates cut certified evaluation cost. Across 432 surrogate fits over six task-regime conditions, the audit statistic tracks deployed search performance at Spearman rank correlation 0.80-0.99, while the rank correlation of R^2 with deployed regret falls as low as 0.33; audited screening reduces certified oracle cost by a measured factor of 25.

Comments47 pages (18 pages main text, 29 pages Supporting Information), 9 figures, 2 tables, 70 references. SI contains complete proofs of all theorems, extended experiments, and the verification methodology

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑