3DS:基于分解难度的数据选择实现LLM的医学领域自适应
3DS: Medical Domain Adaptation of LLMs via Decomposed Difficulty-based Data Selection
- School of Computer Science, Peking University, Beijing, China(北京大学计算机学院)
- Key Laboratory of High Confidence Software Technologies, Ministry of Education(教育部高可信软件技术重点实验室)
- National Engineering Research Center for Software Engineering, Peking University, China(软件工程国家级工程研究中心)
- Big Data Technology Research Center, Nanhu Laboratory, Jiaxing, China(南湖实验室大数据技术研究中心)
- Peking University Information Technology Institute, Tianjin Binhai, China(北京大学信息科学技术研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出两阶段以模型为中心的数据选择框架3DS,通过显式对齐过滤冗余数据,并基于指令理解、响应置信度和响应正确性三个指标进行难度分解选择,有效提升了LLMs在医疗领域的自适应性能。
AI中文摘要:
大型语言模型(LLMs)在通用任务中表现出色,但由于缺乏特定领域的知识,在医疗保健等专业领域仍面临困难。用于领域自适应的监督微调(SFT)数据构建通常依赖于启发式方法,如GPT-4标注或人工数据选择,并以数据为中心,侧重于假定多样、高质量的数据集。然而,这些方法忽略了模型固有的知识分布,引入了噪声、冗余和无关数据,导致所选数据与模型的学习任务不匹配,造成次优性能。为了解决这个问题,我们提出了一个以模型为中心的两阶段数据选择框架,即分解难度数据选择(3DS),该框架将数据与模型的知识分布对齐以实现优化的自适应。在第一阶段,我们通过显式对齐应用提示驱动的数据选择,模型基于其内部知识过滤无关或冗余数据。在第二阶段,我们执行分解难度数据选择,数据选择由我们定义的难度分解引导,使用三个指标:指令理解、响应置信度和响应正确性。此外,基于注意力的重要性加权机制捕获token重要性,以实现更准确的难度校准。这种两阶段方法确保所选数据不仅与模型的知识和偏好对齐,而且对模型学习具有适当的挑战性,从而实现更有效和有针对性的领域自适应。在医学领域的案例研究中,我们在真实世界医疗数据集上的大量实验证明了3DS优于现有方法,准确率提升了5.29%以上。我们的数据集和代码已在 https://github.com/PuppyKnightUniversity/3DS 开源。
英文摘要:
Large Language Models(LLMs) excel in general tasks but struggle in specialized domains like healthcare due to limited domain-specific knowledge.Supervised Fine-Tuning(SFT) data construction for domain adaptation often relies on heuristic methods, such as GPT-4 annotation or manual data selection, with a data-centric focus on presumed diverse, high-quality datasets. However, these methods overlook the model's inherent knowledge distribution, introducing noise, redundancy, and irrelevant data, leading to a mismatch between the selected data and the model's learning task, resulting in suboptimal performance. To address this, we propose a two-stage model-centric data selection framework, Decomposed Difficulty Data Selection (3DS), which aligns data with the model's knowledge distribution for optimized adaptation. In Stage1, we apply Prompt-Driven Data Selection via Explicit Alignment, where the the model filters irrelevant or redundant data based on its internal knowledge. In Stage2, we perform Decomposed Difficulty Data Selection, where data selection is guided by our defined difficulty decomposition, using three metrics: Instruction Understanding, Response Confidence, and Response Correctness. Additionally, an attention-based importance weighting mechanism captures token importance for more accurate difficulty calibration. This two-stage approach ensures the selected data is not only aligned with the model's knowledge and preferences but also appropriately challenging for the model to learn, leading to more effective and targeted domain adaptation. In the case study of the medical domain, our extensive experiments on real-world healthcare datasets demonstrate the superiority of 3DS over exisiting methods in accuracy by over 5.29%. Our dataset and code has been open-sourced at https://github.com/PuppyKnightUniversity/3DS.