发表机构
IIIT Naya Raipur(奈拉普尔印度国际信息技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对低资源场景下LLM诊断系统可靠性不足的问题,提出以临床医生为核心的NSIDDx神经符号鉴别诊断框架,实现离线运行并提炼出相关设计原则。
AI 中文摘要
基于大语言模型(LLM)的诊断系统在基准测试中实现了较高的语义准确性,但针对临床不常见表现的开放式评估显示,其 headline 准确率与可验证的临床可靠性之间存在系统性差距。我们在两个队列中评估了 LLM+罕见病检索增强生成(RAG)流水线,结果表明该范式会生成置信度高但常无法验证、且系统性抵抗临床医生质询的输出。我们提出 NSIDDx(神经符号集成鉴别诊断系统),这一设计框架主张低资源场景下的鉴别诊断(DDx)系统必须将临床医生视为主动推理主体。我们通过具备三元症状编码、矛盾检测、审计字符串及临床医生覆盖功能的神经符号流水线实现该框架,该流水线可在消费级硬件上离线运行。我们提炼出临床医生在环的临床自然语言处理的五项设计原则,并邀请开展前瞻性研究以大规模验证该主张。
英文摘要
LLM-based diagnostic systems achieve high semantic accuracy on benchmarks, but open-ended evaluation on clinically uncommon presentations reveals a systematic gap between headline accuracy and verifiable clinical reliability. We evaluate an LLM+rare-disease-RAG pipeline across two cohorts and show that the paradigm produces confident outputs that are frequently unverifiable and systematically resistant to clinician interrogation. We present NSIDDx (Neuro-Symbolic Integrated Differential Diagnosis System), a design framework arguing that DDx systems in low-resource settings must treat the clinician as an active reasoning agent. We instantiate this through a neuro-symbolic pipeline with ternary symptom encoding, contradiction detection, audit strings, and practitioner override - running offline on consumer hardware. We distill five design principles for clinician-in-the-loop clinical NLP and invite the prospective studies needed to validate the claim at scale.
Comments13 pages, 2 figures, Github: https://github.com/joetheguide2/NSIDDX-