OmniMed-Jev:通过系统一校准LVLM置信度以实现可信的医疗多模态决策
OmniMed-Jev: Calibrating LVLM Confidence for Trustworthy Medical Multimodal Decisions via System One
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
OmniMed-Jev通过将医疗决策表示为候选集上的概率分布,校准LVLM置信度,大幅降低校准与可靠性误差,同时保持点预测性能,使报告概率真实反映决策准确性。
AI中文摘要:
医疗模型的评判不仅基于正确性,还基于所报告的置信度是否与实际准确率相匹配。通用多模态医疗模型扩展了单一模型所能感知的范围,但它们仍将诊断、发现或细胞计数等有界决策表示为生成的文本,因此所报告的概率反映的是下一个词元而非决策本身。受Jev等决策原生界面的启发,我们提出了OmniMed-Jev,它将每个医疗决策表示为对运行时提供的候选集合上的选择(Choice)、否定(Noul)或评分(Score)决策,并返回该集合上的完整分布:互斥类别、发现的二元存在性或一个有界的有序值。该设计在三个方面具有全模态(omni)特性:它接受多样化的成像模态,覆盖不同的预测任务,并通过一个候选条件概率模型来表达它们,从而使异构输出变为可比较的概率而非特定于任务的字符串。在与基于相同主干、数据和训练计划训练生成式基线进行的界面受控比较中,OmniMed-Jev所报告的概率对观测正确性的追踪要紧密得多,将校准误差降低最多达一个数量级,可靠性误差降低最多达两倍,同时点预测性能保持相当;计数是生成式基线仍保持领先的唯一任务族。将决策分布作为模型的输出并非格式上的改变,而是将所报告的数字转变为名副其实的概率的关键。这些结果支持将显式决策建模作为在所评估任务内使所报告置信度具有意义的一种方式,但它们并不构成临床就绪性的证据:该比较无法将界面与相关的训练差异区分开来,我们已随结果一并说明这一点。代码可在以下网址获取:此http URL。
英文摘要:
Medical models are judged not only on correctness, but on whether reported confidence matches actual accuracy. Generalist multimodal medical models have expanded what a single model can perceive, yet they still express bounded decisions such as diagnoses, findings or cell counts as generated text, so the reported probability reflects the next token rather than the decision itself. Motivated by decision-native interfaces such as Jev, we introduce OmniMed-Jev, which represents each medical decision as a Choice, Noul or Score decision over a runtime-supplied candidate set and returns a full distribution over that set: mutually exclusive classes, binary presence of a finding, or a bounded ordered value. The design is omni in three respects: it accepts diverse imaging modalities, covers different prediction tasks, and expresses them through one candidate-conditioned probability model, so heterogeneous outputs become comparable probabilities rather than task-specific strings. In an interface-controlled comparison against a generative baseline trained on the same backbone, data and schedule, OmniMed-Jev's reported probabilities track observed correctness far more closely, reducing calibration error by up to an order of magnitude and reliability error by up to two, while point-prediction performance remains comparable; counting is the one family where the generative baseline stays ahead. Making the decision distribution the model's output is not a format change but what turns reported numbers into probabilities that mean what they say. These results support explicit decision modeling as a way to make reported confidence meaningful within the evaluated tasks, and they are not evidence of clinical readiness: the comparison cannot separate the interface from associated training differences, which we state alongside the results. Code is available at github.com/lytang63/OmniMed-Jev.