arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAGE:语义锚引导演化用于基础医学问答数据合成

SAGE: Semantic Anchor-Guided Evolution for Grounded Medical QA Data Synthesis

Chuan Li, Chengyu Wang, Cen Chen, Ye Lyu, Mingyuan Fan, Ming Gao

arXiv 2610.08093首次发表:更新:

发表机构

East China Normal University; Alibaba Group(华东师范大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对医学问答训练数据稀缺问题,提出SAGE框架,利用语义锚引导小型本地模型迭代合成高质量数据,无需外部API,实验证明其优于传统方法,提升数据效率。

AI 中文摘要

开发用于临床任务(如医学问答(QA))的可靠模型,严重受到高质量、专家标注训练数据有限性的制约。这一挑战因严格的隐私要求以及在资源有限的临床环境中利用大型开源语料库或专有云API的不切实际性而加剧。为解决这些障碍,我们引入了SAGE(语义锚引导演化),一种新颖的数据合成框架,使小型本地部署模型能够生成高质量的医学训练数据。SAGE利用轻量级、公开可用的分类法(如MeSH)作为语义锚,施加结构化先验,有效引导和基础数据生成过程。其核心在于,SAGE迭代地交错原子(基于单个概念)和关联(基于关系)合成,从最小种子引导训练数据。该方法消除了对大型医学文档集合或外部API依赖的需求,为本地数据创建提供了实用解决方案。在多个医学问答基准上的广泛实验表明,使用SAGE合成数据微调的模型始终优于使用自衍生或传统基于文档范式训练的模型,突显了医学LLM开发中数据效率和资源利用的切实改进。代码可在该https URL获取。

英文摘要

Developing reliable models for clinical tasks, such as Medical Question Answering (QA), is severely constrained by the limited availability of high-quality, expert-annotated training data. This challenge is exacerbated by stringent privacy requirements and the impracticality of utilizing large open-source corpora or proprietary cloud APIs within resource-limited clinical settings. To address these obstacles, we introduce SAGE (\textit{Semantic Anchor-Guided Evolution}), a novel data synthesis framework that enables small, locally deployed models to generate high-quality medical training data. SAGE leverages lightweight, publicly available taxonomies such as MeSH as semantic anchors, imposing a structured prior to effectively guide and ground the data generation process. At its core, SAGE iteratively interleaves atomic (individual concept-based) and associative (relation-based) synthesis, bootstrapping training data from minimal seeds. This approach eliminates the need for large collections of medical documents or reliance on external APIs, providing a practical solution for on-premises data creation. Extensive experiments across multiple medical question-answering benchmarks demonstrate that models fine-tuned with SAGE-synthesized data consistently outperform those trained using self-derived or conventional document-based paradigms, highlighting tangible improvements in data efficiency and resource utilization for medical LLM development. Code is available at https://github.com/DIaacKr/SAGE.

CommentsEMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑