发表机构
Indian Institute of Technology Kharagpur; Florida International University(印度理工学院卡拉格普尔分校; 佛罗里达国际大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态临床记录中跨患者组件错误链接的安全问题,提出轻量级预融合筛查模型SHIFT-M3,通过LLM生成与临床报告摘要的对齐一致性检测,在78万条记录上实现高精度异常识别。
AI 中文摘要
多模态临床人工智能通常假设记录中附带的波形、报告、元数据和下游预测属于同一患者。在实践中,链接失败可能静默地组装出各自合理但跨患者的组件,从而造成标准预测模型无法检测的安全问题。我们将此问题作为多模态记录完整性分诊进行研究:给定一个组装好的记录,其模态是否应被信任为属于同一整体?我们引入了SHIFT-M3,一种轻量级的基于文本的预融合筛查方法,用于衡量两个独立生成的ECG文本视图之间的对齐一致性:一个由LLM生成的解释和一个临床报告摘要。在784,680条MEETI ECG记录上,SHIFT-M3在全文本视图交换中实现了97.6%的TPR@5% FPR(AUROC 0.996),在部分交换中实现了90.3%(AUROC 0.974),在标签匹配的困难负样本中实现了97.7%(AUROC 0.996),且仅使用573,569个参数。与同数据集的词汇基线相比,在部分交换和困难负样本上的增益最大,表明模型学习的不仅仅是表面重叠。我们还引入了CMST(冲突类型多模态压力测试)评估分类法、三种子稳定性研究、损失消融、时间容忍度扫描和共享令牌掩码控制。主要的剩余失败模式是纵向模糊性:在默认操作点,同一患者的跨就诊对仍产生87.0%的II型假阳性。
英文摘要
Multimodal clinical AI typically assumes that the waveform, report, metadata, and downstream predictions attached to a record belong to the same patient. In practice, linkage failures can silently assemble individually plausible but cross-patient components, creating a safety problem that standard predictive models are not designed to detect. We study this problem as multimodal record integrity triage: given an assembled record, should its modalities be trusted to belong together? We introduce SHIFT-M3, a lightweight text-based pre-fusion screen that measures alignment-based consistency between two separately produced ECG text views: an LLM-generated interpretation and a clinical report summary. On 784,680 MEETI ECG records, SHIFT-M3 achieves 97.6% TPR@5% FPR for full text-view swaps (AUROC 0.996), 90.3% for partial swaps (AUROC 0.974), and 97.7% for label-matched hard negatives (AUROC 0.996) with only 573,569 parameters. Compared with same-dataset lexical baselines, the gains are largest on partial swaps and hard negatives, suggesting that the model is learning more than surface overlap. We also introduce the CMST (Conflict-type Multimodal Stress Test) evaluation taxonomy, a three-seed stability study, a loss ablation, a temporal-tolerance sweep, and a shared-token masking control. The main remaining failure mode is longitudinal ambiguity: at the default operating point, same-patient cross-visit pairs still produce 87.0% Type-II false positives.
CommentsThis paper has been accepted and presented at MLHC 2026. Please cite from the official proceedings