arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14551cs.CLcs.DLcs.IRcs.LG

用于大语言模型辅助系统评价筛选的辅助不确定性信号:八项Cohen药物分类综述的基准测试

Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews

Arya Rahgozar, Pouria Mortezaagha

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM辅助系统评价筛选的不确定性问题,提出BERT+GCN辅助分类器的不确定性信号,在8个Cohen药物分类数据集上验证了提示策略的效率,发现仅弃权路由为帕累托最优且LLM无法自我分类。

中文摘要 AI 辅助

大语言模型(LLM)正越来越多地用于系统评价中的标题-摘要筛选,但它们的决策缺乏校准的不确定性。我们表明,辅助的BERT+GCN分类器可提供结构化的不确定性信号,提升LLM筛选效率,且我们确定了能最大化投入产出比的提示传递策略。我们使用3个随机种子×5折分层交叉验证(共600个折级结果),在Cohen(2006)基准的8个药物分类数据集上评估5种LLM提示传递条件。每个折训练的BERT+GCN模型通过两项谱检验(代数根检验和范畴悖论检验)将每份测试论文分类为“纳入”“排除”或“弃权(不执行)”。这些条件在信息内容(无/标签/全分数)、选择性(所有论文vs仅弃权论文)和时机(主动vs被动双轮)上存在差异。针对3个数据集的gpt-4.1-mini跨模型试点测试了跨代迁移效果。三项发现:(i)全上下文传递在F1值(+0.011,配对Wilcoxon检验p=0.008)和WSS@95值(+0.050,p=0.039)上有显著提升,同时保持召回率,仅带来1.28倍的令牌成本溢价;(ii)仅弃权路由是帕累托最优的:在仅1.05倍基线成本下达到最高平均召回率(0.92)和AUC-ROC(0.54)——仅为全上下文开销的六分之一;(iii)双轮设计将22.2%±8.8%的记录升级,但从未修改其决策(所有数据集和折的翻转率为0%),这提供了决定性证据,表明当前指令微调的LLM无法自我分类。跨模型试点显示,两种LLM代际的召回率提升均为0.8%。对20796份观察结果的逐论文消融实验表明,双悖论检验在经验上可简化为单行logit差距准则。我们发布了完整流程;从缓存的LLM响应出发,600次运行的实验可在1小时内复现。

英文摘要

Large language models (LLMs) are increasingly used for title-abstract screening in systematic reviews, but their decisions lack calibrated uncertainty. We show that an auxiliary BERT+GCN classifier supplies a structured uncertainty signal that improves LLM screening efficiency, and we identify the prompt-delivery strategy that maximises the benefit-to-cost ratio. We evaluate five LLM prompt-delivery conditions on eight drug-class datasets from the Cohen (2006) benchmark using 3 seeds x 5-fold stratified cross-validation (600 fold-level results). A BERT+GCN model trained per fold classifies each test paper as INCLUDE, EXCLUDE, or MAYBE via two spectral tests (algebraic radical and categorical paradox). Conditions vary information content (none / label / full scores), selectivity (all papers vs. MAYBE only), and timing (proactive vs. reactive two-pass). A cross-model pilot against gpt-4.1-mini on three datasets tests cross-generation transfer. Three findings: (i) Full-context delivery yields significant gains in F1 (+0.011, paired Wilcoxon p=0.008) and WSS@95 (+0.050, p=0.039) at a 1.28x token-cost premium, while preserving recall. (ii) MAYBE-only routing is Pareto-optimal: highest mean recall (0.92) and AUC-ROC (0.54) at only 1.05x baseline cost -- one sixth of full-context overhead. (iii) The two-pass design escalates 22.2% +/- 8.8% of records yet never revises its decision (0% flip rate across all datasets and folds), giving decisive evidence that current instruction-tuned LLMs cannot self-triage. The cross-model pilot shows an identical +0.8% recall uplift for both LLM generations. A per-paper ablation across 20,796 observations shows the dual paradox test reduces empirically to a one-line logit-gap criterion. We release the full pipeline; the 600-run experiment replays in under one hour from cached LLM responses.

发表机构

  • University of Ottawa(渥太华大学)
  • Ottawa Hospital Research Institute(渥太华医院研究所)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑