发表机构
Confidential Core AI(机密核心人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究智能体验证器的自决性问题,发现 LLMs 无法可靠区分可固定检查与需裁判的规则,提出 CoVer 方法通过佐证后验证来构建边界,而非依赖模型自决。
AI 中文摘要
智能体的验证器面临两类规则:一类可以通过固定检查解决,另一类需要裁判。一个自行制定检查的团队会预先确定这种划分。当需求来自外部(如金融、医疗和法律领域)时,智能体执行的是并非其自行编写的规则,因此这种划分在运行时才确定,并且对每个规则的每个谓词、每个动作都要重复进行,其频率之高令任何审查者都无法审计。每种升级方案都假设模型能够自行做出该决定,即具有自决性。在包括欧盟《人工智能法案》、FINRA 指南和已部署的信贷智能体在内的六个语料库中,我们收集了由三个实验室构建的四个模型产生的大约 22,000 个标签。在答案显而易见的地方,它们几乎完全一致,但在监管文本上却出现分歧;它们的错误方向相反,因此没有哪个模型可被信赖为保守选择;并且在已部署智能体自身的规则集上,它们会共同出错,过度声称固定检查即可解决问题,而这一方向永远不会被升级。我们引入了 CoVer(先佐证后验证),它将一致同意视为提名,仅当为谓词合成的检查在干预下仍能通过时才接受该谓词,干预包括读取智能体无法编写的字段以及在确定性改写下保持成立。该门控拒绝了大部分被佐证错误接受的内容,其代价是覆盖率下降,我们对此进行了报告而非调整消除。显而易见的替代方案——与参考裁判达成一致——并不能证明任何东西:它在校准区间内从 30% 攀升至 77%,而真正可决的部分并未变化,因为从被指控群体中抽取的裁判会认可其共有的盲点。自决性并非可从模型中引出的能力,而是验证器必须构建的边界。
英文摘要
A verifier for an agent faces rules of two kinds: the ones a fixed check can settle and the ones that require a judge. A team that derives its own checks fixes that split up front. Where the requirements come from outside, as in finance, healthcare and law, the agent enforces rules it did not write, so the split falls to runtime, recurring for every predicate of every rule on every action at a rate no reviewer can audit. Every escalation scheme assumes a model can make that decision itself, that it is self-decidable. Across six corpora, including the EU AI Act, FINRA guidance and a deployed credit agent, we collect roughly 22,000 labels from four models built by three labs. They agree almost perfectly where the answer is obvious and collapse on regulatory text; their errors run in opposite directions, so no model can be trusted as the conservative choice; and on the deployed agent's own rule-set they err together, over-claiming that a fixed check will do, the direction that never gets escalated. We introduce CoVer (corroborate-then-verify), which treats unanimity as a nomination, admitting a predicate only when the check synthesized for it survives intervention, reading fields the agent cannot write and holding under deterministic rewording. That gate rejects most of what corroboration wrongly admits, at a cost in coverage we report rather than tune away. The obvious alternative, agreement with a reference judge, certifies nothing: it climbs from 30% to 77% across calibration bands while the genuinely decidable share does not move, because a judge drawn from the population under indictment ratifies the blind spot it shares. Self-decidability is not a capability to elicit from a model but a boundary the verifier must construct.
CommentsAccepted at the Who Verifies the Agents? Workshop at NeurIPS 2026. Workshop papers are non-archival. 21 pages (9 content pages plus references and appendices)