评估设计影响专家标注与自动标注的医学主题词差距:Cohen基准上词袋模型与BiomedBERT的对照比较
Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark
浏览论文内容
中文总结 AI 辅助
研究在Cohen基准上比较词袋模型与BiomedBERT作为分类器特征时,专家标注与自动标注的医学主题词差距。通过不同评估设计,发现结果受设计影响大,如他汀类药物在不同设计下差距不同,且存在表示不对称,结论会因评估设计变化而改变。
中文摘要 AI 辅助
系统综述始于有人阅读数千篇摘要以识别少数相关摘要,分类器用于对阅读进行优先级排序。其输入通常会用医学主题词(MeSH)增强,MeSH要么由专家索引员在出版数周或数月后分配,要么由自动工具立即分配。据我们所知,尚未将二者作为分类器特征直接比较,也没有先前的工作探讨这种比较的结果是否取决于分类器的评估方式。我们使用Cohen等人(2006年)关于三个主题的药物类别基准,对词袋逻辑回归分类器(七次重新运行)和BiomedBERT(五个种子)进行特征描述,然后研究他汀类药物的结果在替代设计下如何变化。在标准的5折全语料库设计下,他汀类药物的词袋专家与自动标注差距为+0.096 WSS@95%。将语料库大小与较小主题匹配(n = 803)可将其降至+0.033(95%自举置信区间包含零),全尺寸下的10折交叉验证降至+0.021(置信区间勉强排除零)。在标准评估下,BiomedBERT给出+0.020,在词袋10折结果的抽样噪声范围内。功效分析表明,在阿片类药物或多动症方差下,他汀类药物大小的效应无法检测到,因此那些零结果是设计限制而非信息性的。存在表示不对称:当附加专家MeSH术语时,15.1%的他汀类药物输入超过BiomedBERT的512词元限制,因此截断可能导致变压器差距较小,尽管在此无法与训练量分开。在使用变压器或10折词袋的筛选管道中,测试主题上的差距约为0.02 WSS@95%,置信区间至少在一个边界上跨越零。更广泛地说,关于特征来源的基准结论在评估设计的合理变化下可能会发生重大变化。
英文摘要
A systematic review begins with someone reading thousands of abstracts to identify the few that are relevant, and classifiers are used to prioritise that reading. Their inputs are often augmented with Medical Subject Headings (MeSH), assigned either by expert indexers weeks or months after publication or by automatic tools at once. We did not identify prior work comparing the two directly as classifier features, or asking whether that comparison's outcome depends on how the classifier is evaluated. Using the Cohen et al. (2006) drug-class benchmark, we compare expert assignment against one mechanical procedure, substring matching against a MeSH vocabulary drawn from the benchmark, across a bag-of-words logistic regression classifier (seven reruns) and BiomedBERT (five seeds) on three topics. Under the canonical 5-fold full-corpus design the bag-of-words gap on Statins is +0.096 WSS@95%. Stratified subsampling to matched corpus size (n=803) reduces it by roughly two thirds, to +0.033, with a bootstrap interval that includes zero; 10-fold cross-validation at full corpus size reduces it by roughly four fifths, to +0.021. BiomedBERT under canonical evaluation gives +0.020, a difference of 0.001 from the bag-of-words 10-fold result. An empirical power analysis on a single canonical run per topic indicates that a Statins-sized effect at the per-fold variances of the other two topics would not have been detectable at that design (MDE 0.254 for Opioids, 0.384 for ADHD); at the pooled fold count of the multi-run protocol the bound depends on an effective sample size the design does not determine. The results bound the specific lexical matcher tested rather than automatic MeSH indexing in general. More broadly, benchmark conclusions about feature sources can change substantially under reasonable changes to the evaluation design.