arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00366cs.LGcs.AI

反事实脆弱性证书:结构化证据失效下的高置信度脆弱性暴露

Counterfactual Fragility Certificates: Exposing High-Confidence Brittleness under Structured Evidence Failure

Filippo Cenacchi, Longbing Cao, Runze Yang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出反事实脆弱性证书(CFC),用于审计表格决策系统中高置信度预测的脆弱性,CFC-FDS在识别脆弱高置信度案例上表现优于现有方法,为模型可靠性评估提供了新框架。

中文摘要 AI 辅助

高测试准确率与良好的整体校准无法表明单个预测是否得到其证据的结构化支持。在表格决策系统中,当某一特征族不可用、延迟、含噪、过时或可信度低,而模型仍保持高置信度时,往往会出现故障。现有的校准、不确定性、选择性预测、解释及扰动方法提供标量分数或归因图,但无法提供可重新计算的审计对象,以回答:在声明的证据失效协议下,何种轨迹会使该预测失去支持?我们提出反事实脆弱性证书(Counterfactual Fragility Certificates,CFC),这是一种与模型无关的协议级审计证书(而非形式化鲁棒性证书),它将每个预测映射为有序的证据失效轨迹,该轨迹由贪心翻转预算、归一化边际崩溃面积、退化阈值及脆弱性主导性分数汇总。在7个表格基准及强大的线性、基于树、提升及神经基线模型上,CFC-FDS以0.915的AUROC识别出独立的脆弱高置信度案例,较最强的非证书分数提升了0.405。该优势在扰动、置换重要性、组-SHAP、基线选择、种子方差、预算审查及自然领域不可用性检查中均保持。在20%的审查预算下,CFC-FDS捕获了88.9%的脆弱高置信度案例,而置信度和能量分数仅为31.8%-37.4%。我们还评估了脆弱性感知正则化和脆弱性感知温度校正作为次要用途。CFC为暴露普通分数中心评估遗漏的高置信度脆弱性提供了具体的可靠性框架。

英文摘要

High test accuracy and good aggregate calibration do not show whether an individual prediction is structurally supported by its evidence. In tabular decision systems, failures often occur when a feature family becomes unavailable, delayed, noisy, stale, or low-trust while the model remains highly confident. Existing calibration, uncertainty, selective-prediction, explanation, and perturbation methods provide scalar scores or attribution maps, but not a recomputable audit object answering: under a declared evidence-failure protocol, what trajectory makes this prediction lose support? We introduce Counterfactual Fragility Certificates (CFC), a model-agnostic protocol-level audit certificate-not a formal robustness certificate-that maps each prediction into an ordered evidence-failure trajectory summarized by greedy flip budget, normalized margin-collapse area, degradation thresholds, and fragility dominance score. Across seven tabular benchmarks and strong linear, tree-based, boosting, and neural baselines, CFC-FDS identifies independently brittle high-confidence cases with 0.915 AUROC, improving over the strongest non-certificate score by +0.405. The advantage persists across perturbation, permutation-importance, group-SHAP, baseline-choice, seed-variance, budgeted-review, and naturalistic field-unavailability checks. Under a 20% review budget, CFC-FDS captures 88.9% of brittle high-confidence cases, compared with 31.8-37.4% for confidence and energy scores. We also evaluate fragility-aware regularization and brittleness-aware temperature correction as secondary uses. CFC provides a concrete reliability framework for exposing high-confidence brittleness missed by ordinary score-centric evaluation.

发表机构

  • Macquarie University(麦考瑞大学)

机构由 AI 辅助整理,请以论文原文为准。

↑