ARISMA:AI与大语言模型辅助的系统评价、范围评价与映射研究指南
ARISMA: Guidelines for AI- and LLM-Assisted Systematic Reviews, Scoping Reviews, and Mapping Studies
浏览论文内容
中文总结 AI 辅助
本文提出ARISMA标准,明确AI辅助评价的适用场景、验证方式及人工决策要求,构建可审计的AI辅助证据合成实践指南。
中文摘要 AI 辅助
随着搜索量、更新周期及合成要求持续扩大,系统评价、范围评价、映射研究及相关证据合成工作采用完全手动流程开展的难度日益增加。与此同时,人工智能、机器学习及大语言模型正快速融入评价实践,应用于查询制定、筛选、提取、分类、评价支持及报告等环节。然而,相关实证证据仍不均衡、依赖具体任务,且不足以支持无约束自动化。PRISMA 2020、PRISMA-S、PRISMA-ScR、PRISMA-P、PRESS及SWiM等现有标准仍至关重要,但均未提供端到端操作标准,明确AI应用在方法学上何时适用、应如何验证、哪些评价决策必须由人主导,以及应如何报告AI参与情况以便读者审计。本文提出ARISMA(AI Reporting and Integration standard for Systematic Methods and Analysis,系统方法与分析的AI报告与整合标准),将AI视为经检查、基准测试、记录且可逆的助手,而非自主评价者,其核心治理原则为:每项重要科学决策必须保持可人工解释、可人工审计及人工负责。本文贡献包括生命周期分类法、流程指导、评价全流程各步骤建议、治理与溯源模型、工具支持框架、AI整合报告清单及验证矩阵,还涵盖法律、隐私、基础设施及可持续性考量。该框架通过结构化专家咨询迭代完善,最终形成一套实用且可审计的负责任AI辅助证据合成指南。
英文摘要
Systematic reviews, scoping reviews, mapping studies, and related evidence syntheses are increasingly difficult to conduct with fully manual workflows as search volumes, update cycles, and synthesis requirements continue to expand. At the same time, artificial intelligence, machine learning, and large language models are rapidly entering review practice across query formulation, screening, extraction, categorization, appraisal support, and reporting. Yet the empirical evidence remains uneven, task-dependent, and insufficient to justify unconstrained automation. Existing standards such as PRISMA 2020, PRISMA-S, PRISMA-ScR, PRISMA-P, PRESS, and SWiM remain essential, but none provides an end-to-end operational standard for when AI use is methodologically appropriate, how it should be validated, which review decisions must remain human-led, and how AI involvement should be reported so that readers can audit it. This paper proposes ARISMA, an AI Reporting and Integration standard for Systematic Methods and Analysis. ARISMA treats AI as an inspected, benchmarked, logged, and reversible assistant rather than an autonomous reviewer. It is built around one governing principle: every consequential scientific decision must remain human-interpretable, human-auditable, and human-accountable. The paper contributes a lifecycle taxonomy, process guidance, stepwise recommendations across the review pipeline, a governance and provenance model, a tool-support framework, an AI-integrated reporting checklist, and a validation matrix. It also addresses legal, privacy, infrastructure, and sustainability considerations. The framework was iteratively refined through structured expert consultation. The result is a practical and auditable guideline for responsible AI-assisted evidence synthesis.
发表机构
- University of Southern Denmark(南丹麦大学)
机构由 AI 辅助整理,请以论文原文为准。