arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20938cs.AI

无理由不判断:面向版本化AI评估器的反事实收据

No Judgment Without a Reason: Counterfactual Receipts for Versioned AI Evaluators

  • Alibaba Group(阿里巴巴集团)
  • Cheung Kong Graduate School of Business(长江商学院)

机构由 AI 辅助整理,请以论文原文为准。

Ye Chen, Weining Zhang

AI总结:

该研究针对AI评估器的推理问责问题,提出反事实收据方法与ReasonBench基准,发现模型标准准确率掩盖鲁棒性缺陷,需解耦预测与认证以实现可信审计。

AI中文摘要:

在智能体系统中,评估器常通过有缺陷的推理生成正确标签,这是一个关键缺陷,这类系统负责控制行动、路由审核或提供训练反馈。标准评估仅验证最终标签的正确性,忽略判断变化是否源于有效证据、一致规则或规则的恰当适用性。我们通过三个核心来源(依据、规范和权威)将评估器的推理问责制形式化,改变这些来源可生成一个八单元的反事实判断立方体,用于表征判断更新。我们将判断收据定义为最小的来源替换集,该集合可复现修订后的裁决以解释判断转变。我们推导了黑盒评估器的认证成本边界,并提出了ReasonBench,这是一个包含可验证收据的策略与逻辑推理基准,涵盖19520个案例和7200个对照样本。在冻结评估中,Qwen3-1.7B达到98.41%的收据准确率,而立方体预测得分为96.99%,这一1.42分的一致降幅已通过Qwen3-0.6B的复现得到验证。较高的标准准确率掩盖了严重的鲁棒性缺陷,保留语义的来源置换将有效收据恢复率降至直接预测的54.8%和立方体预测的49.2%。在简单单源变化上训练的模型保留了93.75%的裁决准确率,但仅恢复了复杂多源更新的7.16%的收据。置换再训练将一致性提升至96.6%,却加剧了立方体预测的缺陷。结构化反事实监督无法保证鲁棒推理,我们表明,感知理由的评估必须将预测与认证解耦,报告转换一致性及标准准确率,以实现可信的评估器审计。

英文摘要:

Evaluators often produce correct labels via flawed reasoning, a critical failure for agentic systems gating actions, routing reviews, or supplying training feedback. Standard evaluation only verifies final label correctness, ignoring whether judgment changes stem from valid evidence, consistent rules, or proper rule applicability. We formalize evaluator reasoning accountability via three core sources: grounds, norms, and authority. Varying these sources yields an eight-cell counterfactual judgment cube to characterize judgment updates. We define judgment receipts as minimal source replacement sets that reproduce revised verdicts to explain judgment transitions. We derive certification cost bounds for black-box evaluators and present ReasonBench, a policy and logical reasoning benchmark with verifiable receipts covering 19,520 cases and 7,200 controls. In frozen evaluations, Qwen3-1.7B reaches 98.41% receipt accuracy, while cube prediction scores 96.99%, a consistent 1.42-point drop validated by Qwen3-0.6B replication. Strong standard accuracy masks severe robustness flaws. Meaning-preserving source permutations reduce valid receipt recovery to 54.8% and 49.2% for direct and cube prediction. Models trained on simple single-source changes retain 93.75% verdict accuracy but recover only 7.16% of receipts for complex multi-source updates. Permutation retraining boosts consistency to 96.6% yet worsens cube prediction deficits. Structured counterfactual supervision fails to guarantee robust reasoning. We show reason-aware evaluation must decouple prediction and certification, reporting transformation consistency alongside standard accuracy for trustworthy evaluator auditing.

补充信息

↑