用大型语言模型检测临床试验中的“Spin”
Detecting Spin in Clinical Trials with Large Language Models
- University of Ljubljana(卢布尔雅那大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究针对临床试验中占比超50%的“Spin”问题,开发基于LLMs的自动结局切换检测系统,经2496样本测试,F1达0.78、准确率0.90,优于基线模型,还生成并评估了分类解释。
AI中文摘要:
临床试验中的“Spin”包括扭曲结果呈现的报告实践,在医学领域尤为关键——超过50%未达到统计学显著性的随机对照试验存在“Spin”。比较主要结局与报告结局对检测包括结局切换在内的多种“Spin”至关重要。我们使用300对标注了语义相似度的结局,开发了自动检测结局切换的系统;利用生成的相似度得分和约登指数(Youden index)评估基线文本相似度模型与开源大型语言模型(LLMs),以确定分类阈值。所提方法涉及提示工程、基于token概率的分类及最终决策的多数投票。在含2496个样本的测试集上,该方法的F1分数为0.78、准确率为0.90,优于基线文本相似度模型,但落后于BERT的微调版本;我们还利用LLMs为分类实例生成自然语言解释并人工评估其质量。
英文摘要:
Spin in clinical trials includes reporting practices that distort the presentation of results. This is particularly critical in medicine, where spin is present in more than 50% of randomized controlled trials that fail to reach statistical significance. The comparison of primary and reported outcomes is crucial for detecting several types of spin, including outcome switching. We used 300 pairs of outcomes labeled with semantic similarity to develop a system for automatic detection of outcome switching. We evaluated baseline text similarity models and open-source LLMs using generated similarity scores and the Youden index to determine the classification threshold. The proposed approach involves prompt engineering, classification based on token probabilities, and majority voting for the final decision. The results on the test set of 2,496 examples with an F1 score of 0.78 and an accuracy of 0.90 outperform baseline text similarity models but trail behind fine-tuned versions of BERT. We used LLMs to generate natural language explanations for the classified instances and manually assessed their quality.