发表机构
Nanjing Medical University; The First Affiliated Hospital of Nanjing Medical University; The Friendship Hospital of Ili Kazakh Autonomous Prefecture; Nanjing University(南京医科大学; 南京医科大学第一附属医院; 伊犁哈萨克自治州友谊医院; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种基于临床指南的质量感知框架,将KG构建时的三元组质量信号融合后传播到问答的子图检索与证据提示,可减少知识遗漏与冲突输出,提升临床问答性能。
AI 中文摘要
大型语言模型已助力从临床指南中构建知识图谱(KG),但提取的三元组在结构有效性和证据支持方面存在差异。与此同时,图增强问答(QA)系统通常在检索过程中优化查询相关性,对KG构建期间产生的质量信息的复用有限,这导致构建时的质量控制与推理时的证据使用之间存在脱节。本研究探究构建时的三元组质量是否可作为下游证据选择与呈现的持久信号。我们提出一种质量感知框架,将结构一致性(SchemaConf)与证据支持(EvidScore)建模为互补维度,并将其融合为每个三元组的质量信号Q(t)。该框架并非仅将质量用于过滤,而是保留Q(t)及衍生的质量层级作为图属性,并将其传播到质量加权子图检索和层级条件证据提示中,同时保留段落级来源。对中国糖尿病临床指南的实验表明,质量信号的效用具有分布依赖性。在跨版本和跨模型偏移下,融合后的Q(t)比任一单独成分提供更强的三元组质量区分度(AUC为0.748,EvidScore为0.703,SchemaConf为0.645)。在基于指南的QA中,传播构建时的质量将所需知识遗漏从16.3%降至5.3%,冲突输出从16.3%降至2.7%,基于证据的精度达81.6%,无效引用几乎为零。盲法临床医生评分显示,完整框架优于无检索(五分制评分4.68 vs. 4.21)且接近最优条件(4.80),跨生成器实验也呈现一致趋势。
英文摘要
Large language models have facilitated knowledge graph (KG) construction from clinical guidelines, but extracted triples vary in structural validity and evidential support. Meanwhile, graph-augmented question answering (QA) systems typically optimize query relevance during retrieval, with limited reuse of quality information produced during KG construction. This creates a disconnect between construction-time quality control and inference-time evidence use. We investigate whether construction-time triple quality can serve as a persistent signal for downstream evidence selection and presentation. We propose a quality-aware framework that models structural conformance (SchemaConf) and evidential support (EvidScore) as complementary dimensions and fuses them into a per-triple quality signal, Q(t). Rather than using quality solely for filtering, the framework retains Q(t) and derived quality tiers as graph attributes and propagates them into quality-weighted subgraph retrieval and tier-conditioned evidence prompting, while preserving passage-level provenance. Experiments on Chinese diabetes clinical guidelines show that the utility of the quality signal is distribution dependent. Under cross-version and cross-model shift, the fused Q(t) provides stronger triple-quality discrimination than either component alone (AUC 0.748 vs. 0.703 for EvidScore and 0.645 for SchemaConf). In guideline-grounded QA, propagating construction-time quality reduces required-knowledge omission from 16.3% to 5.3% and conflicting outputs from 16.3% to 2.7%, with an evidence-grounded precision of 81.6% and near-zero invalid citations. Blinded clinician ratings favor the full framework over no retrieval (4.68 vs. 4.21 on a five-point scale) and approach the oracle condition (4.80), while cross-generator experiments show consistent trends.