arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

结合人工智能的大规模质性研究:Socioscope数据管道的架构、管理与运行

Large-Scale Qualitative Research with AI: Infrastructure, Management and Operation of the Socioscope Data Pipeline

Saadi Lahlou, Juan Pablo Caicedo, Shriya Sekhsaria, Valentine Fournand, Paulius Yamin, Helga Nowotny

arXiv 2608.29751首次发表:更新:

发表机构

Paris Institute for Advanced Study; London School of Economics and Political Science; Complexity Science Hub Vienna(巴黎高等研究院; 伦敦政治经济学院; 维也纳复杂科学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究介绍了Socioscope数据管道的架构与运行,其为大规模质性研究收集多国食品系统相关的多媒体数据,利用AI实现规模化分析,相关成果可供其他团队复用改进。

AI 中文摘要

Socioscope项目是大规模质性研究(LSQR)领域的开创性尝试,旨在收集数百个案例的可比较、开放式、多媒体实地数据,并利用人工智能使这些材料能够实现规模化分析。该研究领域为食品系统,所记录的实体包括食品系统中运作的各类组织:农场、加工商、分销商、零售商、餐馆;以及中观层面塑造其环境的行为主体,如市政当局、政府项目、银行、非政府组织和大学。本文提供了关于如何构建和管理生成的语料库以实现人工智能增强分析的技术参考,详细描述了端到端的数据管道:系统性抽样框架、用于捕捉食品系统内每个举措关系的交易网格、旨在维持参与访谈者访问权限的社会契约、从勘探到访谈的运营链(包括访谈的上传、转录、翻译、质量控制和整理)、来源规则(原始内容不可变,每次转换都有记录),以及设备、人员和流程的配置(包括伦理与GDPR合规)。在第一阶段(2023-2026年),该管道已从31个国家生成686个已记录案例,约1430小时的录音、约45万次语音轮次,以及1260万词的转录文本。本文报告了相关成本、指标、经验教训和局限性,以便其他团队能够复用、适配和改进Socioscope方法。

英文摘要

The Socioscope project is a pioneering effort in Large-Scale Qualitative Research (LSQR) collecting comparable, open-ended, multimedia field data on hundreds of cases and using AI to make the material analysable at scale. The domain studied is the food system. The entities documented are the organisations that act in it: farms, processors, distributors, retailers, restaurants; and, at meso level, the actors that shape their environment, such as municipalities, government programmes, banks, NGOs and universities. This paper provides the technical reference for how the resulting data Corpus was built and managed to enable AI-augmented analysis. It describes the data pipeline end to end: the systemic sampling frame; the transaction grid used to capture each initiative's relations within the food system; the social contract that rewards participating interviewees, aiming to sustain access; the operational chain from scouting to interviews, including their uploading, transcription, translation, quality control and curation; the provenance rules (originals are immutable, every transformation is logged); and the installation of equipment, personnel and processes, including ethics and GDPR compliance. In its first phase (2023-2026) the pipeline produced 686 documented cases from 31 countries: some 1,430 hours of recordings, about 450,000 speech turns, and 12.6 million words of transcript. We report costs, metrics, lessons learned and limitations, so that other teams can reuse, adapt, and improve the Socioscope methodology.

Comments80 pages, 6 figures, 3 tables, 7 appendices. The Data Collection Protocol (V8, 55 pages) is available from the authors on request; its table of contents is reproduced in Appendix 7

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑