arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Kurate:可扩展的科学质量分析

Kurate: Scalable Scientific Quality Analysis

Matthew J. Vowels, Jamie Cummins

arXiv 2610.07306首次发表:更新:

发表机构

Kivira Health; University of Bern(Kivira Health; 伯尔尼大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Kurate系统利用大型语言模型评估科学论文质量,从8个维度对4,347篇论文评分,与专家注释高度一致,实现了可扩展的元科学质量分析。

AI 中文摘要

科学检索系统能够找到与问题相关的论文,但通常不会评估这些论文所提供的证据的质量。我们提出了Kurate,一个利用大型语言模型(LLMs)来评估已发表研究质量的系统。Kurate同时使用论文及其相关文档(例如,研究的试验注册和方案),并将其每项判断与作为判断依据的文本段落相关联。我们将Kurate应用于包含4,347篇论文的语料库(其中3,913篇报告了随机试验),并从8个研究设计与报告维度对每篇论文进行评分,具体包括:统计功效、因果识别、预注册、选择性报告、测量效度、分析预设定、报告透明度以及利益冲突与资助。在整个语料库中,我们发现论文最常出现的问题涉及统计功效、选择性报告和分析预设定,尽管不同临床领域的平均质量存在差异。与60份保留的临床试验文档的专家注释相比,Kurate提取的信息在242个方案评分点中的221个和370个结果发表评分点中的294个与专家标签匹配,AC1值分别为0.94和0.81。以一个声誉卓著的高质量临床试验作为工作示例,我们展示了单篇论文的总体评分如何分解为各个独立的判断,每个判断都与来自试验注册、方案和已发表报告的具体证据相关联。这些结果共同表明,此类大规模质量评估是可行的,并且可用于解决元科学研究问题。

英文摘要

Scientific search systems can find papers that are relevant to a question, but they generally do not assess the quality of the evidence that those papers provide. We present Kurate, a system that uses large language models (LLMs) to assess the quality of published studies. Kurate uses both the paper and its related documents (e.g., the study's trial registration and protocol), and links each of its judgments to the passage of text on which that judgment is based. We applied Kurate to a corpus of 4,347 papers (3,913 of which report randomized trials) and scored each paper on 8 dimensions of study design and reporting: specifically, statistical power, causal identification, preregistration, selective reporting, measurement validity, analysis prespecification, reporting transparency, and conflict of interest and funding. Across the corpus, we found that papers most often exhibited issues with statistical power, selective reporting, and analysis prespecification, although average quality differed between clinical areas. When compared against expert annotations of 60 held-out clinical-trial documents, the information Kurate extracted matched the expert label in 221/242 protocol scorepoints and 294/370 results-publication scorepoints, with AC1 0.94 and 0.81, respectively. Using a well-reputed, high quality clinical trial as a worked example, we show how a single paper's overall grade breaks down into separate judgments, with each linked to specific evidence from the trial's registration, protocol, and published report. Together, these results show that large-scale quality assessment of this kind is feasible, and that it can be used to address meta-scientific research questions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑