arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

生成式人工智能能否自动化元分析的数据提取?以间作研究为例

Can Generative AI Automate Data Extraction for Meta-Analysis? A Case Study on Intercropping Research

Zehao Lu, Xingguo Xiong, Wopke van der Werf, Thijs L. van der Plas, Ioannis N. Athanasiadis

arXiv 2609.35089首次发表:更新:

发表机构

Wageningen University & Research; Zhejiang Academy of Agricultural Sciences(瓦赫宁根大学与研究中心; 浙江省农业科学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估三种基于大语言模型的方法在间作文献元分析数据提取中的表现,发现直接零样本提示效果最佳,但所有方法均未达到完全准确。

AI 中文摘要

元分析是对来自多个来源的信息进行综合,以得出总体结论的方法。农业研究中对于元分析有大量需求,用以综合已知信息并分析总体模式。然而,从已发表文献中提取数据是劳动密集型、耗时且乏味的工作,并且受到研究设计、测量单位和术语缺乏标准化的阻碍。这些挑战在作物物种混合(也称为间作)领域尤为突出。随着大语言模型能力的不断增强,许多近期尝试集中于构建系统和工具以自动化数据收集,但针对人工标注的真实数据进行严格评估往往缺失。在本研究中,我们评估了三种基于大语言模型的方法——直接零样本提示、分阶段工作流和多智能体系统——使用六个开放权重模型从间作文献中提取数据。结果与人工整理的真实数据进行比较,并通过下游统计分析进行评估。总体而言,直接零样本提示是最强且最一致的方法,实现了最高的平均相似度调整F1分数0.577,尽管没有任何一种方法接近完全准确。在下游分析中,大多数模型与方法的组合能够恢复预测变量与结果变量之间关系的方向,但无法准确估计其幅度。

英文摘要

Meta-analysis is the synthesis of information from multiple sources to arrive at an overarching conclusion. There is a large need for meta-analysis in agricultural research to synthesize what is known and analyze overarching patterns. Extracting data from published literature is, however, labor-intensive, time-consuming, and tedious, and is impeded by a lack of standardization in research design, units of measurement, and terminology. These challenges are particularly evident in the domain of crop species mixtures, also called intercropping. With the growing capabilities of LLMs, many recent attempts have focused on building systems and tools to automate data collection, yet rigorous assessment against human-labeled ground truth is often missing. In this research, we evaluate three LLM-based approaches---direct zero-shot prompting, a staged workflow, and a multi-agent system---with six open-weight models to extract data from the intercropping literature. The results are evaluated against the manually curated ground truth and through a downstream statistical analysis. Overall, direct zero-shot prompting is the strongest and most consistent approach, achieving the highest mean similarity-adjusted F1 of 0.577, although none of the approaches is close to fully accurate. In the downstream analysis, most model--approach combinations recover the direction of the relationship between the predictor and outcome variables, but do not estimate its magnitude accurately.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑