发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对AI研究代理结果可信度问题,提出发现认证协议(DCP),通过可执行恢复与反馈测试、核心否决及有限样本界,在SQLite优化和虚拟催化剂控制中实现零恢复,提供统一证据语言。
AI 中文摘要
AI研究代理结合先验知识、公开来源和实验反馈来产生有用的结果。发现认证协议(DCP)将这些结果的主张转化为可执行的恢复和反馈测试。第一关在密封评估上验证有用的改进。第二关向匹配的代理提供注册的起始信息和观察到的Web内容,同时隐藏目标研究历史。每个达到数值目标的有效方法都会提供恢复见证并触发核心否决。DCP核心要求充分的对照、零次观察到的恢复,以及在一个新的注册情节中对恢复的有限样本界。可选的第三关测量从共享检查点出发,相对于指定中性策略的真实反馈的平均效果。DCP证据在独立零校准和注册效果边际之后添加此效果。两项受控审计在SQLite优化和虚拟催化剂控制下使用不同模型完整执行了该协议。每项在96个情节中产生零恢复,上限为0.0468。每项配对研究产生30次真实恢复和零次中性恢复,并通过60对零研究。额外案例行使核心、恢复和审计不完整的决策。一个确定性的、无LLM的验证器从冻结证据中重现决策。DCP为AI研究中的有用结果、替代路径和反馈效果提供了共同的证据语言。
英文摘要
AI research agents combine public information and experimental feedback to produce measurable results. The Discovery Certification Protocol (DCP) turns an outcome claim into an executable audit under a registered model, information boundary, and budget. Gate 1 validates useful improvement. Gate 2 tests recovery by matched agents given the starting information and observed Web content, with run history and new measurements withheld. Core requires adequate registered controls, zero recoveries, and a finite-sample recovery bound. Optional Gate 3 compares truthful and neutral feedback from a shared checkpoint; Evidence adds a supported effect and a null-policy equivalence check. Controlled SQLite and virtual catalyst audits pass both decision kernels. On real-data response surfaces, Yacht and Ionosphere pass the Core kernel after zero recoveries in 96 attempts, with an upper bound of 0.0468. Each target combines ten observed utilities and six predictions into a 16-entry data product. Yacht scores 0.7677 on reconstruction of all 32 switch effects, with utility-prediction MAE 0.0315 on its six unmeasured configurations. Fresh truthful continuations recover the target level in 9/30 and 16/30 trials, respectively, separating achieved utility from process repeatability. A deterministic verifier reproduces these local decisions from frozen records.