arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19025cs.AIcs.DB

自提示与跨模型共识:利用大语言模型从科学文献中实现可复现的数据提取

Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models

Valentin Romanov, Monique Bax, Steven Niederer

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出结合自提示与跨模型共识的方法,通过四个工作流程优化LLMs从科学文献提取数据的性能,为科学数据规模化整理提供可行方案。

中文摘要 AI 辅助

从研究论文中准确提取细微、语境化的数据既费力又耗时。本研究调查了前沿的基于浏览器的大语言模型(LLMs)在提取高度语境化信息方面的性能。我们展示了四个逐步升级的工作流程:1)在给定专家精心设计的提示和研究论文的情况下,大多数前沿LLMs在数据提取方面表现良好,但在解释科学语境和细微差别时可能存在困难;2)在给出简单指令后,LLMs可以自行生成提示,其效果几乎与专家编写的提示相当;3)自主发现研究文献存在困难,智能体要么遗漏参考文献,要么生成幻觉式的参考文献;4)LLMs可以根据已发表的指南创建新数据集,这些数据集与人类专家的判断非常接近,但仍然需要人类参与。这些发现共同明确了一种可审计的分工:专家指定证据标准,模型交叉核对多次提取结果,研究人员解决有争议的案例,为在不放弃专家监督的情况下扩大科学数据整理规模提供了可行途径。

英文摘要

Accurately extracting nuanced, contextualized data from research articles is laborious and time intensive. Here, we investigate the performance of frontier, browser-based large language models (LLMs) to extract highly contextualized information. We demonstrate four escalating workflows, 1) given an expert curated prompt and research articles, most frontier LLMs perform well at data extraction, however can struggle with interpreting scientific context and nuance, 2) given simple instructions, LLMs can author their own prompts which were almost as eNective as expert-written prompts, 3) autonomous discovery of research literature was diNicult, agents either missed or hallucinated references, and 4) LLMs can create new datasets from published guidelines that closely match human-expert judges, but still require a human-in-the-loop. Together, these findings define an auditable division of labour in which experts specify the evidence standard, models cross-check repeated extractions and researchers resolve disputed cases, providing a practical route to scaling scientific data curation without relinquishing expert oversight.

↑