发表机构
Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对生物医学文献发表偏倚导致的证据漂移问题,提出DACG-agent智能体,通过KL散度监视和Bradley-Terry奖励模型的两层停止策略,在减少67%检索步骤的同时将证据漂移从15.7%降至6.4%,并提升总体准确性。
AI 中文摘要
自动化生物医学证据综合依赖于检索已发表的研究,但生物医学文献系统性地偏向于阳性发现。因此,更深入的检索可能使系统在真实效应为零时更可能错误地推断出获益。我们将这一现象形式化为“证据漂移”,并证明在标准发表偏倚模型下,零效应查询的假阳性概率随检索深度呈严格递增的大样本包络线,趋近于1。在由140个Cochrane衍生查询组成的留出测试集上,当检索预算从3步增加到20步时,漂移从7.9%单调上升至15.7%,且集中于零效应类别。我们提出DACG-agent,一种漂移感知的因果图智能体,它从PubMed摘要中增量构建因果知识图谱,并应用具有互补作用的两层停止策略:KL散度监视器检测后验收敛(准确性层),以及Bradley-Terry过程奖励模型(PRM),其在线下降检测在证据质量达到峰值时停止检索(效率层)。与全预算检索相比,DACG-agent将证据漂移从15.7%降至6.4%,零效应准确性提高21个百分点(40.0%→61.4%),同时使用减少67%的检索步骤;总体准确性从61.4%升至69.3%(95%置信区间61-77)。模拟证实漂移结果从所分析的投票计数聚合器转移到部署的噪声OR聚合器。
英文摘要
Automated biomedical evidence synthesis depends on retrieving published studies, but the biomedical literature is systematically skewed toward positive findings. Deeper retrieval can therefore make a system \emph{more} likely to falsely infer benefit when the true effect is null. We formalise this phenomenon as \emph{evidence drift} and prove that, under a standard publication-bias model, the false-positive probability on null-effect queries follows a strictly increasing large-sample envelope in retrieval depth, approaching one. Empirically, on a held-out test set of 140 Cochrane-derived queries, drift rises monotonically from 7.9\% to 15.7\% as the retrieval budget grows from 3 to 20 steps, and concentrates in the null-effect class. We present DACG-agent, a drift-aware causal-graph agent that incrementally builds a causal knowledge graph from PubMed abstracts and applies a two-layer stopping policy with complementary roles: a KL-divergence monitor that detects posterior convergence (the accuracy layer), and a Bradley--Terry process reward model (PRM) whose online decline detection halts retrieval once evidence quality peaks (the efficiency layer). Against full-budget retrieval, DACG-agent reduces evidence drift from 15.7\% to 6.4\% and improves null-effect accuracy by 21 percentage points (40.0\%$\to$61.4\%) while using 67\% fewer retrieval steps; overall accuracy rises from 61.4\% to 69.3\% (95\% CI 61--77). A simulation confirms the drift result transfers from the analysed vote-counting aggregator to the deployed noisy-OR one.