发表机构
University of Cambridge; University of Wisconsin–Madison(剑桥大学; 威斯康星大学麦迪逊分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CrossAudit是面向智能体科学的原生Git跨供应商审计协议,通过不同供应商智能体按人类规则审计工作,试验显示不同供应商对规则解读有差异,其自身审计记录为核心证据。
AI 中文摘要
AI科学家不应给自己的作业打分。然而在我们研究的系统中,审核工作的智能体通常与生成该工作的智能体来自同一模型家族,或至少来自同一供应商。已知模型评估器会偏向自身生成的内容。训练方式相似的模型是否也存在共同的盲区,这一问题仍属推测,尚未形成定论,但如果确实存在,审核者便会继承作者的盲区。被标记内容与被放行内容的记录通常存储在外部人员无法重放的平台日志中。我们提出CrossAudit,一种用于监督自主研究流程的协议,该协议基于三项核心约定:每项工作增量均由来自不同供应商的智能体,依据人类编写并进行版本控制的规则手册进行审计;报告、裁决、争议及裁定均以Git提交形式存储,因此监督历史可被重读和引用,不过原始模型交互暂未纳入该记录;任何模型运行前都会执行脚本化检查, advisory judgement( advisory judgement:咨询性判断)不会阻塞流程,模型仅能通过引用规则来阻止流程,且任何模型均不得免除确定性失败。在有限轮次修订后仍存在的阻塞项会提交给人工处理。我们将该协议表述为八项不变量,描述了基于GitHub Actions及数百行Python代码构建的参考实现,并报告了一款密切相关变体在计算化学流程中的实际部署情况。我们还开展了带种子缺陷的试验(共30个工作增量、43个种子缺陷,每种配置各运行一次),对我们自身仓库的跨供应商审计随后揭示了其存在的盲区,我们采纳该审计的发现并报告修正后的结果。该试验显示,两家供应商对同一规则手册的解读存在差异,但未表明其中任何一家更优。此处最有力的证据是针对本论文本身的跨供应商审计所形成的、不受控制的已提交记录。
英文摘要
An AI scientist should not grade its own homework. Yet in the systems we examined, the agent that reviews the work usually comes from the same model family as the agent that produced it, or at least from the same vendor. Model evaluators are known to favour their own generations. Whether models trained alike also share blind spots is a conjecture, not a settled finding, but if they do, the reviewer inherits the author's. The record of what was flagged and what was waved through often sits in platform logs that nobody outside can replay. We present CrossAudit, a protocol for supervising autonomous research pipelines. It rests on three commitments. Each increment of work is audited by an agent from a different vendor against a rulebook a human wrote and versioned. Reports, verdicts, disputes and rulings are git commits, so the supervision history can be re-read and cited; raw model exchanges are not yet part of that record. Scripted checks run before any model does. Advisory judgement never gates the pipeline: a model blocks only by citing a rule, and no model may waive a deterministic failure. Blockers that survive a bounded number of revision rounds go to a person. We state the protocol as eight invariants. We describe a reference implementation built from GitHub Actions and a few hundred lines of Python, and report a live deployment of a closely related variant in a computational-chemistry pipeline. We also ran a seeded-defect trial (30 increments, 43 seeded defects, one run per configuration). A cross-vendor audit of our own repository then voided its blinding. We adopt that audit's findings and report the corrected results. The trial shows that two vendors read the same rulebook differently. It does not show that either is better. The strongest evidence here is the committed, uncontrolled record of cross-vendor audits of this paper itself.
Comments19 pages, 4 figures, 3 tables, 22 references. Reference implementation, audit ledger, and experiment artefacts: https://github.com/dongzhaohe321418-lab/crossaudit