arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

知识综合评审框架:面向多源证据综合的基于大语言模型系统的任务级基准测试

Knowledge Synthesis Review Framework: Task-Level Benchmarking of LLM-Based Systems for Multi-Source Evidence Synthesis

Wafa Shafqat, Mark Patterson, Steven N. Liss

arXiv 2608.12741首次发表:更新:

AI 中文总结

本研究提出KSR人机协作框架,将证据综合拆分为4项任务,对4款LLM系统开展任务级基准测试,发现各系统各有优劣,该框架可揭示单源综合遗漏的信息,且透明可审计。

AI 中文摘要

快速发展领域的证据分散在学术研究、行业报告、政策文件和媒体来源中,这些来源在质量、结构和用途上存在差异,导致及时综合变得困难。大语言模型(LLM)可能会加速这项工作,但它们在评审的不同认知任务中的可靠性仍不确定。我们引入了知识综合评审(KSR),这是一个人机协作框架,将证据综合分解为筛选、提取、分析和综合四个任务,针对每个任务以专家参考标准为基准测试基于LLM的系统,并在持续的专家验证下将每个任务分配给表现最佳的系统。我们在一个包含四种来源类型的1893篇关于AI与工作的文档语料库中,选取了244篇文档作为基准子集,对GPT-5、Claude Sonnet 4、Gemini 2.5 Pro和NotebookLM进行评估,评估采用具有高评分者间信度的黄金标准(一致性达92.2%,kappa值为0.80)。没有系统在所有任务中都领先:Claude Sonnet 4的筛选准确率最高(82.8%),而GPT-5的召回率最高(91.8%),但特异性较低;提取任务中,标题和来源的一致性超过90%,但作者和参考文献字段的一致性有所下降;在解释性分析和跨来源综合任务中性能下降最明显,专家判断仍然必不可少;对截止后文档的污染检查显示,没有证据表明预先接触会夸大结果。将该框架应用于完整语料库时,分配后的工作流程揭示了单源综合会遗漏的跨来源不对称性和盲区,包括员工福祉、小型企业和全球南方。KSR提供了一个透明、可审计、与模型无关的框架,用于管理研究综合中的LLM辅助,同时保留人类问责制。

英文摘要

Evidence in rapidly evolving fields is fragmented across academic studies, industry reports, policy documents, and media sources that differ in quality, structure, and purpose, making timely synthesis difficult. Large language models (LLMs) may accelerate this work, but their reliability across the distinct cognitive tasks of a review remains uncertain. We introduce the Knowledge Synthesis Review (KSR), a human-in-the-loop framework that decomposes evidence synthesis into screening, extraction, analysis, and synthesis, benchmarks LLM-based systems on each task against expert reference standards, and routes each task to the best-performing system under continuous expert validation. We evaluated GPT-5, Claude Sonnet 4, Gemini 2.5 Pro, and NotebookLM on a 244-document benchmark subset drawn from a 1,893-document corpus on AI and work spanning four source types, against a gold standard with high inter-rater reliability (92.2% agreement, kappa = 0.80). No system led on all tasks. Claude Sonnet 4 achieved the highest screening accuracy (82.8%) and GPT-5 the highest recall (91.8%) at the expense of lower specificity. Extraction exceeded 90% agreement for titles and sources but degraded in author and reference fields. Performance declined most in interpretive analysis and cross-source synthesis, where expert judgment remained essential. A contamination check on post-cutoff documents showed no evidence that prior exposure inflated results. Applied to the full corpus, the routed workflow surfaced cross-source asymmetries and blind spots that single-source synthesis would miss, including worker well-being, small firms, and the Global South. KSR offers a transparent, auditable, model-agnostic framework for governing LLM assistance in research synthesis while preserving human accountability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑