arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TCMQA:一个包含38K问题的中医基准,附有执业医师参考答案

TCMQA: A 38K-Question Traditional Chinese Medicine Benchmark with a Licensed-Practitioner Reference

Tzu-Heng Huang, Jet Lin, Eric Lin

arXiv 2609.33014首次发表:更新:

AI 中文总结

提出TCMQA,一个含38,279道中医执业考试题及101位执业医师15,151条参考答案的开放基准,评估29个模型,发现中文预训练数据比模型规模更能预测中医能力,且模型与医师难度不转移。

AI 中文摘要

语言模型的医学基准几乎完全建立在西方生物医学之上。中医是一个独立的体系,拥有自己的诊断框架和文献,且在很大程度上未被测量。现有的少数中医评估规模小、范围窄,且很少与人类参考答案配对。我们提出了TCMQA,一个包含来自中国中医执业考试的38,279个问题的开放基准,并配有来自101位执业医师的15,151条回答。我们评估了来自9个家族的29个指令微调模型,参数规模从0.27B到14.8B不等。准确率跨度超过59个百分点,且没有模型接近饱和。预训练数据对中医能力的预测远优于模型规模:一个12B的西方预训练模型达到39.6%的准确率,而一个规模仅为其八分之一的中文预训练模型达到60.8%。有9个模型超过了执业医师的多数投票准确率64.9%,其中最佳者高出21.8个百分点,且这9个模型全部来自同一中文预训练家族。然而,难度在模型和执业医师之间并不转移:模型准确率在执业医师评定的难度上保持平坦,所有29个模型的条目级一致性接近零,并且在8.4%的条目上,执业医师正确而领先模型错误。我们在以下https URL发布了语料库、执业医师回答、测试框架和逐条目模型输出。

英文摘要

Medical benchmarks for language models are built almost entirely on Western biomedicine. Traditional Chinese Medicine (TCM) is a separate system, with its own diagnostic framework and its own literature, and it remains largely unmeasured. The few TCM evaluations that exist are small, narrow, and rarely paired with a human reference. We present TCMQA, an open benchmark of 38,279 questions from Chinese TCM licensing examinations, paired with 15,151 responses from 101 licensed practitioners. We evaluate 29 instruction-tuned models from 9 families, spanning 0.27B to 14.8B parameters. Accuracy ranges over 59 points, and no model approaches saturation. Pretraining data predicts TCM ability far better than scale: a 12B Western-pretrained model reaches 39.6%, while a Chinese-pretrained model an eighth its size reaches 60.8%. Nine models exceed the practitioner majority vote of 64.9%, the best by 21.8 points, and all nine come from that same Chinese-pretrained family. Yet difficulty does not transfer between models and practitioners: accuracy is flat across practitioner-rated difficulty, item-level agreement is near zero for all 29 models, and on $8.4\%$ of items the practitioners are correct where the leading model is wrong. We release the corpus, the practitioner responses, the harness, and per-item model outputs at https://huggingface.co/datasets/TechTCM/TCMQA.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑