ClaimReceipt:智能体评估中证据充分性与覆盖范围的验证
ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations
浏览论文内容
中文总结 AI 辅助
提出ClaimReceipt规范与验证器,可验证智能体评估中声明的证据充分性与覆盖范围,在历史记录和前瞻性实验中验证有效,开销极低但规范可读性待提升。
中文摘要 AI 辅助
智能体评估面临两个不同的证据问题:一是报告的声明是否可从保留的证据中重新计算(充分性),二是保留的记录是否覆盖了已提交的实验集(覆盖范围)。通用日志和哈希链接的转录本均无法可靠回答这两个问题。我们提出ClaimReceipt,这是一种与声明相关的收据规范及选择性验证器,它将类型化交易证据绑定到已签名的实验清单,并针对每个声明返回通过(PASS)、无效(INVALID)或不确定(INCONCLUSIVE)结果。我们在实现前冻结了该规范(SHA-256 18d109...b81)。在1392条历史买卖双方记录上,CR-2验证器重现了全部5个人工标注的审计结果,准确重放了600条确定性记录和792条后生成记录,在测试的消融实验中使13个已声明的字段组全部非冗余,且在11/11语义故障上返回预期结果,假阳性为0/8。随后我们运行了独立的前瞻性CR-3轮次:在推理前提交了30个任务,终端收据被签名并链式存储,私有证据为审计员加密。完整证据会使覆盖范围和核算返回通过;扣留1份终端收据返回覆盖范围不确定(INCONCLUSIVE_COVERAGE),而扣留所有私有公开内容则保留覆盖范围和协议验证,但使经济声明不确定,完全符合预先注册的预测。收据检测仅增加了模型推理时间的0.021%和每笔交易9.9KB的开销。一项规范可读性探测显示,我们冻结的规范目前对独立读者而言尚不明确。因此,声明验证既需要与声明充分的证据,也需要一个已提交的、能让遗漏变得可见的整体集合。
英文摘要
Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from retained evidence (sufficiency), and whether the retained records cover the committed experiment set (coverage). Generic logs and hash-linked transcripts answer neither reliably. We introduce ClaimReceipt, a claim-relative receipt specification and selective verifier that binds typed transaction evidence to a signed experiment manifest and returns PASS, INVALID, or INCONCLUSIVE per claim. We freeze the specification before implementation (SHA-256 18d109...b81). On 1,392 historical buyer--seller records, a CR-2 verifier reproduces all five manually labeled audit verdicts, exactly replays 600 deterministic and 792 post-generation records, makes every one of 13 declared field groups non-redundant under tested ablations, and returns the expected result on 11/11 semantic faults with 0/8 false positives. We then run a separate prospective CR-3 epoch: 30 assignments are committed before inference, terminal receipts are signed and chained, and private evidence is encrypted for an auditor. Complete evidence yields coverage and accounting PASS; withholding one terminal receipt returns INCONCLUSIVE_COVERAGE, while withholding all private openings preserves coverage and protocol verification but makes economic claims inconclusive, exactly matching a preregistered prediction. Receipt instrumentation adds 0.021% of model-inference time and 9.9 KB per transaction. A specification-legibility probe indicates that our own frozen specification is not yet unambiguous to an independent reader. Claim verification therefore requires both claim-sufficient evidence and a committed universe against which omissions become visible.
发表机构
- Blossom AI
- Blossom AI Labs(Blossom AI实验室)
机构由 AI 辅助整理,请以论文原文为准。