arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

系统综述LLM筛选中的类别不平衡与批次效应

Class Imbalance and Batch Effects in LLM-Based Screening for Systematic Reviews

Gilberto Sussumu Hida, Danilo Monteiro Ribeiro, Clayton Suguio Hida

arXiv 2608.14737首次发表:更新:

AI 中文总结

本研究针对系统综述的LLM筛选场景,发现批次处理对决策行为的影响大于患病率元数据,建议评估批次处理需兼顾成本与决策行为效应。

AI 中文摘要

本研究以系统综述的研究筛选为应用场景,分析大型语言模型(LLMs)在不平衡二分类任务中的表现。实验在5项系统综述中开展,对比了单独处理与批次处理两种模式,同时考察是否纳入患病率元数据的影响。结果显示,患病率元数据的影响有限,无证据表明其能提升性能;相反,批次处理会引发更大的行为变化,且变化幅度随类别患病率不同而有所差异,聚合层面与项目层面的分析结果并非总是一致。因此,评估批次处理时,不仅应考量成本,还需关注其对决策行为的影响。

英文摘要

This study analyses LLMs in imbalanced binary classification, using study screening in systematic reviews as the application domain. An experiment was conducted in five reviews, comparing individual and batch processing, with and without prevalence metadata. The results indicate a limited influence of the prevalence metadata, with no evidence that it improves performance. In contrast, batch processing produced larger behavioral changes that varied according to the prevalence of the class. The aggregate and item-level analyses did not always coincide. Therefore, batch processing should be evaluated not only in terms of cost, but also in relation to its effects on decision-making behavior.

Comments12 pages, 4 figures. Accepted at ENIAC 2026 (National Meeting on Artificial and Computational Intelligence), part of BRACIS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑