剖析用于合成肿瘤学数据生成的神经符号质量保障机制
Dissecting Neuro-Symbolic Quality Assurance for Synthetic Oncology Data Generation
浏览论文内容
中文总结 AI 辅助
本研究通过三项对照研究剖析神经符号质量保障机制,明确符号门控、本体接地等组件的作用,发现其可提升合成肿瘤学数据临床有效性,且检索增强效果依赖模型。
中文摘要 AI 辅助
利用大语言模型生成的合成临床数据可解决癌症分期研究受限的数据稀缺问题,但肿瘤学领域的幻觉属于类别性危害:一项临床不可能的分期分配会污染所有基于该数据训练的下游模型。神经符号流水线在生成过程中进行验证,但单个质量保障组件的贡献仍不明确。我们报告三项对照研究,分别隔离门控必要性、约束归因和检索条件性,在各适配器条件间保持生成协议、多样性阈值和微调超参数恒定。符号门控强制架构完整性、针对医学系统命名法(SNOMED)的本体覆盖,以及符合美国癌症联合委员会第八版规则的分期逻辑一致性。未门控时,29.9%的记录存在架构故障,20.1%包含临床无效分期。架构验证是核心过滤器:在完全门控语料库中,它拒绝512条记录中的148条,本体接地进一步拒绝24条,分期逻辑验证未拒绝任何记录——唯一产生逻辑违规的生成器已在架构阶段被排除,这使得临床逻辑验证成为生成器条件性保障,而非主导过滤器。检索增强高度依赖模型:它使一个生成器的门控合规性提升12.5个百分点,对第二个生成器无显著影响,导致第三个生成器的输出崩溃。在所有门控配置中,本体密度基本不变,表明符号验证可提升临床有效性,而非词汇丰富度。因此,本研究中符号门控可提升语料库有效性,但未对应增加真实肺癌记录的增益;检索需按模型评估,且不应将本体密度作为语料库质量的替代指标。
英文摘要
Synthetic clinical data generation with large language models addresses the scarcity that limits cancer staging research, but oncology hallucinations are categorically harmful: one clinically impossible staging assignment contaminates every downstream model trained on it. Neuro-symbolic pipelines validate during generation, yet the contribution of individual quality-assurance components remains unclear. We report three controlled studies isolating gate necessity, constraint attribution, and retrieval conditionality, holding generation protocol, diversity thresholds, and fine-tuning hyperparameters constant across adapter conditions. The symbolic gate enforces schema completeness, ontology coverage against the Systematized Nomenclature of Medicine, and staging-logic consistency under American Joint Committee on Cancer eighth-edition rules. Ungated, 29.9% of records contain schema failures and 20.1% contain clinically invalid staging. Schema validation is the load-bearing filter: within the fully gated corpus it rejects 148 of 512 records, ontology grounding a further 24, and staging-logic validation none---the only generator producing logic violations is already excluded on schema, making clinical-logic validation a generator-conditional safeguard rather than the dominant filter. Retrieval augmentation is strongly model-dependent: it improves gate compliance for one generator by 12.5 percentage points, has no measurable effect for a second, and collapses output in a third. Across gated configurations ontology density is largely unchanged, indicating that symbolic validation improves clinical validity rather than vocabulary richness. Symbolic gating therefore buys corpus validity but no commensurate gain on real lung-cancer notes in this study; retrieval should be evaluated per model, and ontology density should not be reported as a proxy for corpus quality.
发表机构
- University of North Texas(北得克萨斯大学)
机构由 AI 辅助整理,请以论文原文为准。