arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22417cs.AIcs.CLcs.HC

用于调查文本分析的大语言模型(LLMs):人类与GPT-5在归纳内容分析上的性能比较

LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis

Leonardo Bergmann, Renata Gheorghiu, Ana Gvritishvili, Alex Mican, Chris Stewart, Topias Tolonen-Weckström

首次发表
浏览论文内容

中文总结 AI 辅助

本研究对比人类与GPT-5.4在903份欧洲博士生调查开放式回答的归纳内容分析中的性能,发现GPT-5.4编码的ARI为0.61,可近似人类编码表现,或可作为归纳质性分析的可扩展支持工具。

中文摘要 AI 辅助

大语言模型(LLMs)正越来越多地被用于支持质性研究中的文本分析,但关于其在归纳内容分析中性能的证据仍然有限。本研究对来自欧洲博士生调查的6个变量、共903份开放式调查回答,开展了人类与基于LLM的归纳编码比较研究。5名人类编码员按照标准化编码方案执行归纳内容分析,而LLM(GPT-5.4)则采用既定提示程序完成相同任务。使用调整兰德指数(ARI)评估人类与LLM输出间的一致性。结果显示,人类与LLM之间存在一致性,编码的ARI值为0.61,主题生成的ARI值为0.54。这些数值接近人类内部编码与主题结果的一致性(ARI=0.68)以及LLM内部的一致性(ARI=0.76)。不同变量间的一致性差异显著,实体内部一致性低始终与实体间一致性低相关,凸显了数据特征和个体表现对可靠性的影响。总体而言,研究结果表明,在该特定场景下,LLMs可在编码层面近似人类编码表现,或可作为归纳质性分析的可扩展支持工具。

英文摘要

Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive content analysis remains limited. This study compares human and LLM-based inductive coding of open-ended survey responses from 903 answers across six variables from a European PhD student survey. Five human coders performed inductive content analysis following a standardized coding scheme, while an LLM (GPT-5.4) conducted the same task using an established prompting procedure. Agreement between human and LLM outputs was assessed using the Adjusted Rand Index (ARI). Results showed an alignment between humans and the LLM, with ARI values of 0.61 for coding and 0.54 for theme generation. These values were close to the internal consistency of coding and theme results within humans (ARI = 0.68) and the LLM (ARI = 0.76). Agreement varied widely across variables, with low within-entity consistency consistently linked to low between-entity agreement, underscoring the role of data characteristics and individual performance in reliability. Overall, the findings suggest that LLMs can approximate human coding in this case-specific setting, particularly at the coding level, and may serve as a scalable support tool for inductive qualitative analysis.

↑