发表机构
School of Health & Wellbeing, University of Glasgow; Shanghai Sixth People’s Hospital, Shanghai Jiao Tong University School of Medicine; School of Life Science and Technology, University of Electronic Science and Technology of China(格拉斯哥大学健康与福祉学院; 上海交通大学医学院附属第六人民医院; 电子科技大学生命科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对罕见疾病尾部的选择性预测,研究发现现有小型开放权重LLM在超罕见疾病上表现受限,前两位边际可提升部分病例选择性准确率,且仅未标记分数无法确定边际切换是否有益。
AI 中文摘要
给定患者的临床发现,诊断系统会对可能的疾病进行排序,并必须决定何时认可其首个预测或推迟至审核。该决策通常通过对最高分数设置阈值来做出。针对排序输出的选择性预测需经过两项检查:首先,排序器必须生成足够多的正确首排序预测以确保目标可行。在按疾病患病率分层的2000条患者记录中,8个小型开放权重LLM在超罕见疾病上的Recall@1最多达到4.6%;在10%覆盖率下,即便对其现有预测进行完美的置信度排序,也无法达到50%的选择性准确率。更准确的模型通过了相同检查,表明该限制是特定领域的。其次,置信度信号必须与所做决策匹配。对于固定候选排序器,前两位的边际会抵消跨候选共享的组件;仅使用表型的Exomiser中,该边际以29.0%的准确率选择10%的病例,而整体准确率为13.3%,同时最高分数无法提供可靠的门控。然而,这种抵消可能会移除检测候选列表是否包含答案所需的信息,SciFact检索和生物医学实体链接证实了这种区别。最后,我们证明仅未标记的分数无法确定切换至边际是否会有帮助。
英文摘要
Given a patient's clinical findings, a diagnostic system ranks possible diseases and must decide when to endorse its first prediction or defer it for review. This decision is usually made by thresholding the top score. Selective prediction over ranked outputs begins with two checks. First, the ranker must produce enough correct top-ranked predictions to make the target feasible. Across 2,000 patient records stratified by disease prevalence, eight small open-weight LLMs achieve at most 4.6% Recall@1 on ultra-rare diseases. At 10% coverage, even a perfect confidence ranking of their existing predictions therefore cannot reach 50% selective accuracy. More accurate models pass the same check, showing that the limit is regime-specific. Second, the confidence signal must match the decision being made. For fixed-candidate rankers, the top-two margin cancels components shared across candidates. On phenotype-only Exomiser, it selects 10% of cases at 29.0% accuracy, compared with 13.3% overall, while the top score provides no reliable gate. Yet that cancellation can remove information needed to detect whether the candidate list contains an answer. SciFact retrieval and biomedical entity linking confirm this distinction. Finally, we prove that unlabelled scores alone cannot determine whether switching to the margin will help.