大声还是沉默?一种用于多模态临床AI中模态级失败分析的可复用框架
Loud or Silent? A Reusable Framework for Per-Modality Failure Analysis in Multimodal Clinical AI
浏览论文内容
中文总结 AI 辅助
该研究提出一种模型无关的多模态临床AI模态级失败分析框架,可复用且仅用部署可观测信号,经植入真实值和MIMIC-IV队列验证,能定位失败模态、区分失败类型,为心脏基础模型部署提供参考。
中文摘要 AI 辅助
多模态临床模型通常在所有模态都存在的情况下以准确率进行评估,但部署时可能会缺失模态;在常规使用心电图(ECG)的场景中,超声心动图(echocardiogram)往往不可用。除了准确率损失的大小,两个关键问题需要解决:哪个模态是导致失败的原因,以及当该模态被移除后,模型是“大声”(可检测到)失败还是“沉默”(未被标记的)失败。这种区分是针对每个样本和每个模态的,与事后特征归因(如SHAP)不同。模型会频繁更换,而用于回答上述问题的评估方法可复用。我们提出一种模型无关的模态失败框架:给定N个模态嵌入、任意掩码感知探针和标签,它会返回每个样本的失败分类、将错误归因于模态的每个模态互补矩阵,以及“大声-沉默”的dropout轮廓,该轮廓可区分可监测的失败与远离决策边界未被标记的沉默失败,且仅使用部署时可观测的信号。我们将其作为小型、经过单元测试的工具包发布,并通过植入的真实值进行验证。在不同随机种子下,它能恢复植入的模态主导性和互补子集,报告每个模态的“大声-沉默”率,并扩展到三模态互补矩阵;由于植入的结构是构造已知的,这验证了对每个样本归因的恢复而非临床性能。我们随后在配对的MIMIC-IV队列上,针对冻结的EchoJEPA和HuBERT-ECG嵌入,实例化该框架以评估LVEF和EF≤40%的HFrEF门,在留出的测试集(n=245)上,移除超声心动图几乎使错误翻倍;限制队列规模的超声心动图与ECG的狭窄重叠本身就是心脏基础模型的部署发现。所有工作可在该https URL获取。
英文摘要
Multimodal clinical models are usually judged on accuracy with every modality present, but deployment removes modalities; an echocardiogram is often unavailable where an ECG is routine. Two questions then matter beyond the size of the accuracy loss: which modality was responsible, and whether the model fails loudly or silently once that modality is dropped. The distinction is per-example and modality-level, and is separate from post-hoc feature attribution (e.g. SHAP). Models are replaced often; the evaluation that answers these questions is reused. We present a model-agnostic modality-failure framework: given N modality embeddings, any mask-aware probe, and labels, it returns a per-example failure taxonomy, a per-modality complementarity matrix that attributes error to modalities, and a loud-vs-silent dropout profile separating monitorable failures from those that pass unflagged far from the decision boundary, using only deployment-observable signals. We release it as a small, unit-tested harness and validate it against planted ground truth. Across seeds it recovers that planted modality dominance and complementary subset, reports per-modality loud-vs-silent rates, and scales to a three-modality complementarity matrix; because the planted structure is known by construction, this validates recovery of per-example attribution rather than clinical performance. We then instantiate the framework on frozen EchoJEPA and HuBERT-ECG embeddings for LVEF and the EF <= 40% HFrEF gate over a paired MIMIC-IV cohort, where on the held-out test split (n = 245) dropping echo nearly doubles error. The narrow echo-to-ECG overlap that bounds cohort size is itself a deployment finding for cardiac foundation models. All of our work can be found at https://github.com/criticaldata/PRIMED-AI.
发表机构
- Massachusetts Institute of Technology(麻省理工学院)
- American International School Vienna(维也纳美国国际学校)
- Neuqua Valley High School(纽夸谷高中)
- North Hollywood High School(北好莱坞高中)
- Hopewell Valley Central High School(霍普韦尔谷中央高中)
- University of California, Berkeley(加州大学伯克利分校)
- Siriraj Hospital(诗里拉吉医院)
- Harvard University(哈佛大学)
- Agency for Science, Technology and Research(新加坡科技研究局)
- Johns Hopkins University(约翰斯·霍普金斯大学)
- University of Geneva(日内瓦大学)
- Poznan University of Technology(波兹南理工大学)
机构由 AI 辅助整理,请以论文原文为准。