递归大语言模型在生物医学问答中的退化:一项跨代研究
Recursive LLM Degradation in Biomedical Question Answering: A Cross-Generation Study
浏览论文内容
中文总结 AI 辅助
本研究通过跨代实验,发现递归合成数据训练在生物医学问答中导致模型性能下降,且大模型下降更明显,但未证实临床幻觉或通用模型崩溃。
中文摘要 AI 辅助
反复使用语言模型自身生成的数据进行训练,可能会形成合成数据反馈回路,使错误和分布偏差被重新引入后续的训练数据集。本文使用PubMedQA和两个Qwen2.5模型规模(0.5B和3B参数)研究了生物医学问答(QA)中的这一过程。研究比较了递归合成数据条件(其中第G(k+1)代模型使用第G(k)代生成的答案进行训练)与人类对照条件(反复使用原始人类训练数据)。研究评估了从G0到G3的四代模型,使用两个随机种子(42和123)以及一个包含1,000个专家标记样本的固定评估集。评估指标包括疾病和化学实体F1、上下文支持率、词汇和语义相似度、答案长度、重复率等。对于两种模型规模和两种随机种子,递归条件在疾病实体F1、化学实体F1、上下文支持率、ROUGE-L和余弦相似度上的下降幅度均大于人类对照条件。在固定的无重复三元组解码约束下,观察到的主要行为变化是答案长度增加,而测量的三元组重复率并未增加。3B模型的差异变化幅度大于0.5B模型。这一差异在疾病F1、上下文支持率、余弦相似度和答案长度方面尤为明显。这些结果表明,在生物医学QA中使用递归合成数据训练会带来领域特定的变化,但并未确立临床幻觉率或通用模型崩溃。
英文摘要
Repeatedly training language models on their own generated data may create a synthetic-data feedback loop in which errors and distributional biases are reintroduced into subsequent training datasets. This paper studies that process in biomedical question answering (QA) using PubMedQA and two Qwen2.5 model sizes, 0.5B and 3B parameters. The study compares a recursive synthetic-data condition, in which generation G(k+1) is trained on answers produced by G(k), against a Human-Control condition that repeatedly uses the original human training data. The study evaluates across four generations from G0-G3 with two random seeds (42 and 123) and a fixed evaluation set of 1,000 expert-labeled samples. The evaluation includes disease and chemical entity F1, context-supported rate, lexical and semantic similarity, answer length, repetition rate, and other evaluation metrics. The Recursive condition for both model sizes and both seeds showed larger declines than the Human-Control condition in disease entity F1, chemical entity F1, context-supported rate, ROUGE-L, and cosine similarity. Under the fixed no-repeat 3-gram decoding constraint, the main observed behavioral change was increased answer length, while the measured 3-gram repetition rate did not increase. The magnitude of the difference-in-change was larger for the 3B model than for the 0.5B model. This difference was particularly apparent in disease F1, context-supported rate, cosine similarity, and answer length. These results show domain-specific changes associated with using recursive synthetic-data training in biomedical QA, but do not establish clinical hallucination rates or universal model collapse.
发表机构
- CMR Institute of Technology(CMR理工学院)
机构由 AI 辅助整理,请以论文原文为准。