证据约束推理:胶质母细胞瘤影像基因组学中生物医学AI的神经语义验证
Evidence-Bound Reasoning: Neuro-Semantic Verification of Biomedical AI in Glioblastoma Radiogenomics
浏览论文内容
中文总结 AI 辅助
该研究提出神经语义验证框架,将放射组学数据转化为可验证证据,在胶质母细胞瘤影像基因组学中实现独立于预测性能的确定性验证,确保AI解释有据可依。
中文摘要 AI 辅助
背景:生物医学AI能够生成看似合理的解释,却无法可靠地验证每项陈述是否得到患者特定证据的支持。我们开发了一个神经语义验证框架,将放射组学测量转化为可寻址的证据记录和机器可检查的声明。方法:UPenn-GBM放射组学与独立多中心队列中基于标准化MRI和专家验证分割的de novo CaPTk提取结果对齐。共享空间包含来自T1、T1GD、T2和FLAIR MRI的1,728个特征,覆盖三个肿瘤区域。参考定义的语义状态源自611例UPenn病例。我们评估了跨队列可迁移性、模型关联溯源、确定性验证、受控预测退化以及LLM声明提取试点;MGMT预测仅作为迁移压力测试。结果:语义状态一致性的中位数为0.786(加权kappa为0.709),范围从形态学特征的0.918到强度特征的0.252。外部证据账本包含331名患者的1,655条模型关联记录。验证器在6,620项声明损坏基准中实现了100%的精确集合准确率。在24例试点中,GPT-5.6 Sol复现了72/72个预设原子声明,冻结的验证器恢复了24/24个预期条件。在受控退化过程中,ROC AUC从0.899降至0.500,而验证准确率保持1.000。外部MGMT判别能力较弱(ROC AUC为0.543)。结论:可验证性可以独立于预测性能进行工程设计和评估。LLM可能用于构建解释,而最终的证据一致性检查保持确定性。
英文摘要
Background: Biomedical AI can generate plausible explanations without reliably verifying whether each statement is supported by patient-specific evidence. We developed a neuro-semantic verification framework that converts radiomic measurements into addressable evidence records and machine-checkable claims. Methods: UPenn-GBM radiomics were aligned with de novo CaPTk extraction from standardized MRI and expert-validated segmentations in an independent multicenter cohort. The shared space comprised 1,728 features from T1, T1GD, T2, and FLAIR MRI across three tumor regions. Reference-defined semantic states were derived from 611 UPenn cases. We evaluated cross-cohort transportability, model-linked provenance, deterministic verification, controlled predictive degradation, and an LLM claim-extraction pilot; MGMT prediction served only as a transport stress test. Results: Median semantic-state agreement was 0.786 (weighted kappa 0.709), ranging from 0.918 for morphologic to 0.252 for intensity features. The external evidence ledger contained 1,655 model-linked records for 331 patients. The verifier achieved 100% exact-set accuracy in a 6,620-claim corruption benchmark. In a 24-case pilot, GPT-5.6 Sol reproduced 72/72 prespecified atomic claims, and the frozen verifier recovered 24/24 expected conditions. During controlled degradation, ROC AUC declined from 0.899 to 0.500 while verification accuracy remained 1.000. External MGMT discrimination was weak (ROC AUC 0.543). Conclusions: Verifiability can be engineered and evaluated independently of predictive performance. LLMs may structure explanations, while final evidence-consistency checking remains deterministic.
发表机构
- Faculty of Mathematics and Informatics – Sofia University St. Kliment Ohridski(圣克莱门特奥赫里德斯基索非亚大学数学与信息学院)
机构由 AI 辅助整理,请以论文原文为准。