arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ConfTriage:用于肺结节恶性程度分级的校准感知大语言模型分级框架,结合选择性专家模型转诊机制

ConfTriage: A Calibration-Aware LLM Triage Framework for Pulmonary Nodule Malignancy with Selective Specialist Deferral

Md Rabiul Islam, Samir Abdaljalil, Erchin Serpedin, Hasan Kurban

arXiv 2608.10885首次发表:更新:

发表机构

Texas A&M University; Hamad Bin Khalifa University(德克萨斯农工大学; 哈马德·本·哈利法大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出ConfTriage框架,以校准感知LLM为核心结合专家DL后备模型,通过肺结节属性的自然语言表述实现分级,在LIDC-IDRI数据集上取得88.22%的F1分数与0.92的AUC,可高效处理多数病例并仅转诊不确定病例,为医疗决策支持提供了新路径。

AI 中文摘要

肺结节恶性程度预测通常依赖于经图像训练的专家级深度学习(DL)模型,这类模型需要大量带标注的成像数据及针对特定任务的训练。本研究探究仅通过标准结节属性的忠实自然语言表述进行读取的通用型大语言模型(LLM),能否作为具备校准能力的分级层。我们提出ConfTriage,这是一种基于三大支柱构建的校准方法:以语言作为模态、以校准作为安全机制、为低置信度病例配备选择性专家DL后备模型。我们证明了两项理论保障:一是有限样本联合误差界,可生成明确的每阈值操作证书;二是预言不等式,表明相对于贝叶斯最优转诊分类器的超额风险由LLM概率的L1校准误差控制。在LIDC-IDRI数据集上对五种前沿LLM开展的受控七类输入消融实验显示,自然语言描述主导诊断信号,而低级图像统计量基本无诊断价值。ConfTriage的F1分数达88.22%,AUC为0.92,仅通过零样本LLM推理即可解决76.5%的病例,仅将不确定病例转诊至专家DL后备模型。这些结果表明,结构化放射学描述可捕捉具有临床意义的诊断信息,并由校准后的LLM加以利用以实现选择性转诊。该框架为在医疗决策支持系统中结合通用型LLM预测与专家AI模型提供了可行路径。源代码公开于此https URL。

英文摘要

Pulmonary nodule malignancy prediction typically depends on image-trained specialist deep learning (DL) models that require substantial annotated imaging data and task-specific training. We investigate whether a generalist large language model (LLM), reading only a faithful natural-language rendering of standard nodule attributes, can serve as a calibrated triage layer. We propose ConfTriage, a confidence-calibrated method built on three pillars: language as the modality, calibration as the safety mechanism, and a selective specialist DL backstop for low-confidence cases. We prove two guarantees: a finite-sample combined-error bound yielding an explicit per-threshold operational certificate, and an oracle inequality showing that excess risk over the Bayes-optimal deferral classifier is controlled by the L1 calibration error of the LLM probability. A controlled seven-way input ablation across five frontier LLMs on LIDC-IDRI shows that natural-language descriptions dominate the diagnostic signal, while low-level image statistics are essentially diagnostically vacuous. ConfTriage achieved an F1 score of 88.22% and an AUC of 0.92, resolving 76.5% of cases using zero-shot LLM inference alone and referring only uncertain cases to the specialist DL backstop. These results demonstrate that clinically meaningful diagnostic information can be captured through structured radiological descriptions and leveraged by calibrated LLMs for selective referral. The framework suggests a practical pathway for combining generalist LLM prediction with specialist AI models in medical decision-support systems. Source code is publicly available at https://github.com/rabiul-ai/ConfTriage.

Comments14 pages, 8 figures. Currently under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑