探索用于生物医学数据到文本生成的小语言模型的训练后对齐:以药品说明书为例
Exploring Post-Training Alignment of Small Language Models for Biomedical Data-to-Text Generation: A Case Study of Medication Leaflet
浏览论文内容
中文总结 AI 辅助
该研究对生物医学数据到文本生成任务训练小语言模型的方法进行比较分析,在药品说明书数据集上用基于Qwen的SLMs探索多种训练后方法,通过整理数据评估模型,结果显示对齐后的SLMs性能优异,GRPO跨数据集性能最稳健。
中文摘要 AI 辅助
将复杂的生物医学数据转化为患者易懂的叙述是现代生物医学信息学的核心。本研究对在专门的生物医学数据到文本生成任务中训练小语言模型(SLMs)进行了比较分析。我们在药品说明书数据集上,使用基于Qwen的SLMs探索了广泛采用的训练后方法,包括监督微调(SFT)、直接偏好优化(DPO)、优势比偏好优化(ORPO)和群体相对策略优化(GRPO)。为评估跨数据集的通用性,我们还整理了来自openFDA的药品标签数据。我们使用ROUGE等标准词汇重叠指标以及语义相似性度量来评估模型。实验结果表明:(1)对齐后的SLMs优于GPT-5等专有模型;(2)ORPO优于SFT基线;(3)GRPO在测试对齐方法以及GPT-5中产生了最稳健的跨数据集性能。
英文摘要
Translating complex biomedical data into patient-friendly narratives is central to modern biomedical informatics. This study presents a comparative analysis of training small language models (SLMs) in specialized biomedical datato-text generation tasks. We explore widely adopted post-training methods including supervised fine-tuning (SFT), direct preference optimization (DPO), odds ratio preference optimization (ORPO), and group relative policy optimization (GRPO) with Qwen-based SLMs on a medicine package leaflets dataset. To assess cross-dataset generalizability, we also curated drug label data from openFDA. We evaluate models using both standard lexical overlap metrics like ROUGE as well as semantic similarity measures. Across our experiments, the results show that (1) the aligned SLMs outperform proprietary models like GPT-5; (2) ORPO outperforms the SFTbaselines; (3) GRPO yields the most robust cross-dataset performance among the alignment methods tested as well as GPT-5.