arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20825cs.AI

预测认证无法替代解释认证:复合压力下可信人工智能的能力边界

Prediction certification cannot replace explanation certification: a competence envelope for trustworthy AI under compound stress

Nataliya Shakhovska, Ivan Izonin, Stergios-Aristoteles Mitoulis

AI总结:

研究表明仅预测侧认证无法确保AI可信,提出结合预测与解释认证的能力边界框架,可捕捉预测侧认证遗漏的失效模式,为可信AI认证提供新方法。

AI中文摘要:

人工智能系统越来越多地做出具有重大影响的判断——判断哪位患者病情恶化、哪栋建筑可以安全进入、图像是否真实,并且人们基于其预测的准确性和置信度而信任它们。对这些系统的保障措施相应地基于预测:准确性、校准和保形覆盖均衡量模型的性能。此类检查是否足以确立模型的可信性仍不清楚。本文证明它们无法做到。我们提出一个分离定理,表明可靠模型与受损模型在所有预测侧认证(包括准确性、校准和覆盖)下表现完全相同,但在解释保真度和部署行为上存在任意差异。检测这种失效除了需要模型的预测外,还需要访问模型的决策机制。我们引入能力边界(competence envelope)作为将预测和解释认证结合为单一可部署标准的操作框架。在不同数据集和模型类别上,该框架揭示了仅预测侧认证无法捕捉的失效模式。因此,针对在预测行为中不可见的失效进行认证,需要同时提供关于模型决策机制及其输出的证据。

英文摘要:

Artificial intelligence systems increasingly make consequential judgments - which patient is deteriorating, which building is safe to enter, whether an image is authentic and are trusted on the strength of how accurately and confidently they predict. The safeguards that certify them are correspondingly prediction-based: accuracy, calibration and conformal coverage all measure how well a model performs. Whether such checks are sufficient to establish model trustworthiness has remained unclear. Here we prove that they cannot. We establish a separation theorem showing that a reliable model and a compromised one can be identical under every prediction-side certificate, including accuracy, calibration and coverage, yet differ arbitrarily in explanation fidelity and deployment behaviour. Detecting this failure requires access to the model's decision mechanism in addition to its predictions. We introduce the competence envelope as an operational framework that combines prediction and explanation certification into a single deployable criterion. Across diverse datasets and model classes, the proposed framework reveals failure modes that prediction-side certification alone does not capture. Certification against failures that are invisible in prediction behaviour therefore requires evidence about the model's decision mechanism as well as its outputs.

↑