arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CalTwin:基于Fisher信息正则化的校准、分布偏移鲁棒医学世界模型研究

CalTwin: Towards Calibrated, Shift-Robust Medical World Models via Fisher-Information Regularisation

Behraj Khan, Shabir Ahmad, Syed Ahmad Chan Bukhari, Tahir Qasim Syed

arXiv 2607.26752首次发表:更新:

发表机构

Institute of Business Administration Karachi; Gachon University; St. John’s University(卡拉奇商业管理学院; 嘉泉大学; 圣约翰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出CalTwin方法,结合Fisher信息偏移惩罚与置信度失配惩罚,应用于GRU医学世界模型,在PhysioNet脓毒症挑战赛上显著降低了分布外隐状态预测误差,同时改善了置信度校准。

AI 中文摘要

医学世界模型旨在学习患者或器官生理的隐状态,以及预测该状态在干预下如何演变的转移函数,为从基于影像的诊断到数字孪生治疗规划等下游任务提供支持。这类模型在临床部署中面临两类失效模式:一是协变量偏移,由于训练数据分散在不同医院、扫描仪和时间节点,隐动态预测器在各数据片段中所见的特征分布存在差异,且与部署时的分布不一致;二是置信度失配,多步预测往往在临床风险最高的地方过度自信。我们认为这两个问题可通过一个轻量级正则化目标CalTwin实现统一处理,该方法结合了从我们先前针对碎片化协变量偏移修复的工作中改编的基于Fisher信息的偏移惩罚,以及从我们先前针对校准型视觉-语言分类的工作中改编的置信度失配惩罚,应用于基于GRU的医学世界模型的隐转移预测器。我们推导了该联合目标,明确了从分类场景中直接迁移的证明步骤和需要适配的步骤,并在PhysioNet 2019脓毒症挑战赛上进行评估,将两个医院系统视为连续的训练数据片段,未见过的系统作为分布外测试集。结果显示,与无惩罚基线相比,CalTwin将分布外下一步隐状态的均方误差降低了9.1%(仅Fisher信息矩阵惩罚单独贡献了7.0%);置信度失配惩罚带来的预期校准误差(ECE)降低是切实的但幅度较小(CalTwin为0.7%,仅置信度失配惩罚单独为1.3%)。

英文摘要

Medical world models aim to learn a latent state of patient or organ physiology and a transition function that forecasts how that state evolves under interventions, supporting downstream tasks from imaging-based diagnosis to digital-twin treatment planning. Two failure modes threaten the reliability of such models in clinical deployment: (i)~\emph{covariate shift}, because training data are fragmented across hospitals, scanners, and time, so the feature distribution seen by the latent-dynamics predictor differs across fragments and from the distribution at deployment; and (ii)~\emph{confidence misalignment}, because multi-step forecasts are often overconfident exactly where clinical risk is highest. We argue that both problems admit a unified treatment via a single lightweight regularisation objective, \textbf{CalTwin}, which combines a Fisher-Information-based shift penalty adapted from our prior work on fragmented covariate-shift remediation~\cite{khan2025mitigating,khan2025causal} with a Confidence Misalignment Penalty adapted from our prior work on calibrated vision-language classification~\cite{khan2025confidence}, applied here to a GRU-based medical world model's latent transition predictor. We derive the combined objective, establish which proof steps transfer from the classification setting without modification and which require adaptation, and evaluate it on the PhysioNet 2019 Sepsis Challenge, treating the two hospital systems as sequential training fragments and the unseen system as an out-of-distribution test. CalTwin reduces OOD next-step latent-state MSE by 9.1\% relative to the no-penalty baseline (FIM penalty alone accounts for 7.0\%); the ECE reduction from the Confidence Misalignment Penalty is real but small (0.7\% for CalTwin, 1.3\% for CMP alone).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑