发表机构
Imperial College London(帝国理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对AI生成研究的缺陷验证问题,提出研究契约绑定实验选择等与执行证据,通过实验验证其可检测多数变异案例,为优化下的合规完成度提升提供必要测试方向。
AI 中文摘要
AI生成研究中的部分缺陷可从其产物中识别,另一些则需要执行前已获批准的信息。我们提出研究契约,将声明的实验选择、运行义务和主张范围与记录的执行证据绑定,并区分这种契约相关验证与科学真理。一项使用8组自行编写的干净/变异对的诊断实验阐明了信息边界:确定性检查器将注册的、特定于故障的规则应用于已批准和已执行的对象,检测到全部8个已注册变异。在18个记录的裁判别名获得单独的、经元数据过滤的包(无配对上下文或注册表选定的故障标签)的情况下,144个变异评估案例中有104个收到缺陷标记;剩余案例包括32次弃权(不执行)和8次终端失败,且对变异案例无明确的干净判定。部分包保留了批准和执行字段,包括摘要。提示要求裁判在证据不足时弃权(不执行)。这些结果描述了一个故意信息不对称的开发环境;它们未将权威信息的影响与任务规范和规则选择的差异分离,也不涉及验证器质量或智能体奖励黑客行为。我们确定全信息比较、合法适应控制和闭环智能体评估是必要的测试,以验证契约检查是否在优化下提高了有用的合规完成度。
英文摘要
Some defects in an AI-generated study can be identified from its artifacts; others require knowledge of what was approved before execution. We propose study contracts that bind declared experimental choices, run obligations and claim scope to recorded execution evidence, and distinguish this contract-relative verification from scientific truth. A diagnostic using eight self-authored clean/mutated pairs illustrates the information boundary. A deterministic checker applying a registered, fault-specific rule to approved and executed objects detected all eight registered mutations. Across eighteen recorded judge aliases given individual metadata-filtered packages without pair context or the registry-selected fault label, 104 of 144 mutated evaluation cases received defect flags; the remaining cases comprised 32 abstentions and eight terminal failures, with no explicit clean decisions on mutated cases. Some packages retained approval and execution fields, including digests. The prompt instructed judges to abstain when evidence was insufficient. These results characterize a deliberately information-asymmetric development setting; they do not isolate the effect of authoritative information from differences in task specification and rule selection, and they are not comparative verifier quality or agent reward hacking. We identify full-information comparisons, legitimate-adaptation controls and closed-loop agent evaluations as necessary tests of whether contract checks improve useful compliant completion under optimization.
Comments15 pages, 4 tables. Preprint. Not peer reviewed. Companion manuscripts prepared in parallel