发表机构
University of Virginia; University of Delaware; Memorial Sloan Kettering Cancer Center(弗吉尼亚大学; 特拉华大学; 纪念斯隆凯特琳癌症中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用稀疏自编码器从癌症分割模型内部提取“错误感”信号,通过概念激活特征训练分类器,实现高准确率失败检测并保持分割质量。
AI 中文摘要
癌症分割模型可能无声地失败,生成看似合理但错误的掩膜,从而带来漏诊或非必要活检的风险。一个关键问题随之产生:AI模型是否“知道”自己出错,如果知道,我们能否利用这一信号来预测其自身的失败?人类确实拥有“错误感”(FOE):一种自发的、在思考过程中标记潜在错误的不安感。我们研究癌症分割模型是否表现出类似的内部信号。与输出层面的线索(如预测置信度或不确定性)不同——这些线索无法提供失败原因的任何洞见,并且存在灵敏度与质量之间的权衡,即高检测灵敏度可能降低整体分割质量——我们转而提出从模型的内部工作机理中捕获其“错误感”。利用机制可解释性工具,特别是稀疏自编码器,我们将内部神经激活分解为人类可解释的概念字典,并表明失败案例表现出独特的潜在特征:与成功分割相比,活跃概念更少且激活幅度更低。通过在这些概念激活上训练分类器,我们实现了准确的失败检测,并提供了模型错误的解释。在前列腺、胰腺和脑癌分割上的实验表明,我们的方法在失败检测方面优于基于输出的方法,同时保持了分割质量。
英文摘要
Cancer segmentation models can fail silently, generating plausible but incorrect masks that risk missed findings or unnecessary biopsies. A critical question arises: Do AI models "know" when they are wrong, and if so, can we use the signal to predict their own failures? Humans do have a "Feeling of Error" (FOE): a spontaneous sense of unease that flags a potential error during thinking. We investigate whether cancer segmentation models exhibit an analogous internal signal. Unlike output-level cues (e.g., prediction confidence or uncertainty), which offer no insight into why a failure occurs and suffer from a sensitivity-quality tradeoff where high detection sensitivity could degrade overall segmentation quality. We instead propose to capture the model's FOE from its inner workings. Using mechanistic interpretability tools, specifically Sparse Autoencoders, we decompose internal neural activations into a dictionary of human-interpretable concepts and show that failure cases exhibit a distinct latent signature: fewer active concepts with lower activation magnitudes compared to successful segmentation. By training a classifier on these concept activations, we achieve accurate failure detection along with explanations for the model's mistakes. Experiments on prostate, pancreatic, and brain cancer segmentation demonstrate that our approach outperforms output-based methods in failure detection while preserving segmentation quality.
CommentsIn ECCV 2026