arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24534cs.AI

面向生理安全临床语言模型的神经符号对齐

Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models

  • University of Kurdistan Hewlêr(库尔德斯坦大学)
  • Nanyang Technological University(南洋理工大学)
  • Massachusetts Institute of Technology(麻省理工学院)
  • University of Kragujevac(克拉古耶瓦茨大学)

机构由 AI 辅助整理,请以论文原文为准。

Abdulhady Abas Abdullah, Erik Cambria, Milena Zivkovic

AI总结:

提出神经符号对齐框架,耦合临床LLM与HGNN生理世界模型,在CSB基准上显著提升临床LLM的生理安全性能,优于ORPO、GPT-4等方法,为临床LLM安全对齐提供新路径。

AI中文摘要:

临床大语言模型(LLM)可生成看似符合事实但生理不安全的建议。本研究探究是否能通过将偏好优化建立在结构化生理知识而非纯文本监督之上,来提升安全对齐效果。方法:我们提出神经符号对齐(Neurosymbolic Alignment),这是一种训练时框架,将7B规模的临床LLM与基于HGNN(异构图神经网络)的生理世界模型耦合,该模型依托包含84.7万个节点的生物医学知识图谱。候选响应通过稳态约束、多跳路径合理性及药物相互作用惩罚进行评分,所得排名驱动迭代式在线策略ORPO更新。评估在临床安全基准(Clinical Safety Benchmark, CSB)上开展,该基准包含2500个生成式临床推理中生理约束违反的场景。结果:相较于ORPO,所提方法将CSS从69.5%提升至90.8%(增幅21.3个百分点),在盲法子集上将医师评估的HR从14.1%降至5.1%,并将DID从72.8%提升至91.6%。这些提升得到不依赖HGNN的规则引擎安全评分(RSS:86.4%,较ORPO提升21.2个百分点,与CSS的一致性r=0.97)的佐证。尽管参数规模比GPT-4(5次少样本)小10倍,该方法在所有安全指标上均优于GPT-4,且在CSS上比推理时自校正流水线(SFT+SelfCorrect)高出11.4个百分点。在合成电子病历(EHR)式噪声下,仍能保持84.2%的CSS。消融分析显示,HGNN评分(-16.2个百分点)与迭代训练(-11.5个百分点)是主要贡献因素。针对200个医师标签的PhysioScore校准得到ECE=0.038、kappa=0.91。结论:训练时的生理接地在受控评估下,为开放权重临床LLM带来可测量且可独立验证的安全提升,需在真实临床数据上开展外部验证以确定这些提升是否可迁移至部署场景。

英文摘要:

Clinical LLMs can generate recommendations that are factually plausible yet physiologically unsafe. We investigate whether safety alignment can be improved by grounding preference optimization in structured physiological knowledge rather than text-only supervision. Methods: We propose Neurosymbolic Alignment, a training-time framework that couples a 7B clinical LLM with an HGNN-based Physiological World Model over an 847K-node biomedical knowledge graph. Candidate responses are scored using homeostatic constraints, multi-hop path plausibility, and drug-interaction penalties, and the resulting rankings drive iterative on-policy ORPO updates. Evaluation is performed on the Clinical Safety Benchmark (CSB), a 2,500-scenario benchmark for physiological constraint violations in generative clinical reasoning. Results: Relative to ORPO, the proposed method improves CSS from 69.5% to 90.8% (+21.3 pp), reduces physician-evaluated HR from 14.1% to 5.1% on the blinded subset, and improves DID from 72.8% to 91.6%. These gains are corroborated by an HGNN-independent Rule-Engine Safety Score (RSS: 86.4%, +21.2 pp over ORPO; r=0.97 concordance with CSS). The method also exceeds GPT-4 (5-shot) on all safety metrics despite a 10x parameter disadvantage, and outperforms an inference-time self-correction pipeline (SFT+SelfCorrect) by 11.4 pp CSS. Under synthetic EHR-style noise, 84.2% CSS is retained. Ablation analysis shows that HGNN scoring (-16.2 pp) and iterative training (-11.5 pp) are the dominant contributors. PhysioScore calibration against 200 clinician labels yielded ECE = 0.038 and kappa = 0.91. Conclusion: Training-time physiological grounding produces measurable and independently verifiable safety improvements in open-weight clinical LLMs under controlled evaluation. External validation on real clinical data is required to determine whether these gains transfer to deployment settings

↑