arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

挖掘意义:AI辅助文献综述中的测量误差

Mining Meaning: Measurement Error in AI-Assisted Literature Reviews

Jeffrey D. Michler, Kieran Douglas, Anna Josephson

arXiv 2609.27686首次发表:更新:

发表机构

University of Arizona; University of California - Davis(亚利桑那大学; 加利福尼亚大学戴维斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究将LLM辅助文献综述视为测量问题,通过评估ChatGPT在提取降雨工具变量论文元数据中的表现,发现测量误差对论文级结论影响大,对文献级结论影响小,强调需检验结论对测量系统的稳健性。

AI 中文摘要

研究人员越来越多地使用生成式人工智能,特别是大型语言模型(LLMs),来自动化研究流程中的各项任务。我们研究了这些工具在阅读、分类和综合大量学术文献方面的可靠性。我们将LLM辅助的文献综述视为一个测量问题,将模型视为测量系统,并追踪其误差如何影响下游结论。作为测试案例,我们使用三种不同的ChatGPT实现来识别和提取使用降雨作为工具变量的经济学论文中的元数据。我们将每种实现与一组人工标注的评估数据子集进行基准测试,然后将这些实现部署到全语料库中以提取元数据。LLMs在二元分类任务上表现良好,但随着任务要求更大的上下文解释,性能会下降。更重要的是,研究人员对模型输出的依赖程度不仅取决于阅读任务的复杂性,还取决于数据被要求支持的声明类型。相同程度的测量误差对论文层面的声明有显著影响,而对关于文献的更广泛声明影响甚微。因此,LLM生成数据中的测量误差恰恰在构成LLM相对于人类审稿人主要附加价值的细节层面上影响最大。我们得出结论,标准模型性能指标能提供关于生成数据质量的信息,但本身并不能确立下游推断的可信度。研究人员还必须评估实质性声明是否对生成底层数据的测量系统具有稳健性。

英文摘要

Researchers increasingly use generative AI, particularly large language models (LLMs), to automate tasks across the research pipeline. We study the reliability of these tools at the reading, classification, and synthesis of large bodies of academic literature. We frame LLM-assisted literature reviews as a measurement problem, treating models as measurement systems and tracing how their errors affect downstream conclusions. As a test case, we use three different implementations of ChatGPT to identify and extract metadata from economics papers that use rainfall as an instrumental variable. We benchmark each implementation against a subset of human-labeled evaluation data, and then deploy those implementations to extract metadata from the full corpus. The LLMs perform well on binary classification, but performance deteriorates as tasks demand greater contextual interpretation. More importantly, how much researchers can rely on model outputs depends not only on the complexity of the reading task but also on the type of claims the data is asked to support. The same amount of measurement error substantially affects paper-level claims while having little effect on broader claims about the literature. Measurement error in LLM-generated data is thus most consequential at precisely the level of detail that constitutes an LLM's principal value added over human reviewers. We conclude that standard model performance metrics are informative about the quality of generated data but do not by themselves establish the credibility of downstream inference. Researchers must also evaluate whether substantive claims are robust to the measurement system used to generate the underlying data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑