证候、协同与安全:面向中医处方生成的结构化推理与知识驱动对齐
Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation
浏览论文内容
中文总结 AI 辅助
针对大语言模型在中医处方生成中的推理、调整和安全缺口,提出SFT到K-RL的四阶段框架,显著提升处方质量,7B模型超越GPT-5。
中文摘要 AI 辅助
将大语言模型应用于中医处方生成,揭示了三个临床关键缺口:模型产生端到端的映射,缺乏遵循理法方药范式的可审计推理(SR缺口);孤立地处理每次就诊,不通过随证加减进行后续调整(LA缺口);未能强制执行绝对禁忌规则,如十八反(SC缺口)。我们提出了一个渐进式四阶段框架(SFT → PG-CoT → Dynamic → K-RL),以解决每个缺口:PG-CoT在理法方药范式下约束思维链蒸馏,以产生可审计的诊断链;Dynamic SFT对患者轨迹进行建模,并带有显式的转换推理;K-RL将确定性药理规则编码为基于规则的DPO偏好信号。在12个微调模型和6个零样本基线上,我们的框架显著提高了处方质量,优于零样本基线——其中7B模型(Mistral-7B)在所有三个中医评估指标上超越了零样本GPT-5。
英文摘要
Applying large language models to Traditional Chinese Medicine (TCM) prescription generation reveals three clinically critical gaps: models produce end-to-end mappings without auditable reasoning following the li-fa-fang-yao paradigm (SR Gap), treat each encounter in isolation without follow-up adjustment via sui zheng jia jian (LA Gap), and fail to enforce absolute contraindication rules such as Shi Ba Fan (SC Gap). We propose a progressive four-stage framework (SFT $\to$ PG-CoT $\to$ Dynamic $\to$ K-RL) that addresses each gap: PG-CoT constrains CoT distillation under the li-fa-fang-yao paradigm to produce auditable diagnostic chains, Dynamic SFT models patient trajectories with explicit transition reasoning, and K-RL encodes deterministic pharmacological rules as rule-based DPO preference signals. Across 12 fine-tuned models and 6 zero-shot baselines, our framework substantially improves prescription quality over zero-shot baselines---with a 7B model (Mistral-7B) surpassing zero-shot GPT-5 on all three TCM evaluation metrics.
发表机构
- Tsinghua University(清华大学)
- Guangdong Provincial Laboratory of Traditional Chinese Medicine Hengqin(广东省中医药横琴实验室)
机构由 AI 辅助整理,请以论文原文为准。