发表机构
University of Washington; Colleague AI(华盛顿大学; Colleague AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现以人类共识为标准的一致性评估无法准确衡量LLM定性编码质量,盲专家验证显示部分LLM编码优于人类,提出了可迁移的验证协议与编码分工框架。
AI 中文摘要
对大语言模型(LLM)辅助定性编码的评估几乎普遍将模型性能衡量为与人类编码员的一致性,这种做法假定人类编码是需逼近的标准。本研究提供实证证据表明,该假定存在一致性指标无法检测的失效情况。5个LLM系统和3名经过训练的人类编码员,独立将包含72个条目的分层编码本应用于来自K-12 AI平台的2560条教育者消息。除常规一致性分析外,一名独立领域专家对855组编码集配对比较进行盲评,对人类与机器来源一视同仁。两种评估方法呈现双向分歧:人类-LLM一致性(平均Jaccard系数0.30)远低于人类-人类一致性(0.52),标准实践会据此判定LLM编码更差,但盲验证者对人类与LLM编码的偏好率无差异(51.5%对48.5%,p=0.537),Bradley-Terry排名将2个LLM置于3名人类编码员中的2名之上。针对若干实质性编码,人类共识编码了共享偏差,而验证者更认可LLM的解释。因此,基于一致性的评估不足以支撑自动化决策,本研究还展示了可迁移的验证协议与编码层面的分工框架。
英文摘要
Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes human coding is the standard to approximate. This study provides empirical evidence that the presumption fails in ways agreement metrics cannot detect. Five LLM systems and three trained human coders independently applied a 72-item hierarchical codebook to 2,560 educator messages from a K-12 AI platform. Beyond conventional agreement analysis, an independent domain expert judged 855 pairwise comparisons of code sets blind to source, treating human and machine sources symmetrically. The two evaluation approaches diverge in both directions. Human-LLM agreement (mean Jaccard 0.30) falls well below human-human agreement (0.52), which standard practice would read as inferior LLM coding, yet the blind verifier preferred human and LLM coding at indistinguishable rates (51.5% vs. 48.5%, p = 0.537), and a Bradley-Terry ranking placed two LLMs above two of three human coders. For several substantive codes, human consensus encoded shared bias that the verifier rejected in favor of the LLM interpretation. Agreement-based evaluation is therefore insufficient for automation decisions, and the study demonstrates a transferable verification protocol and a code-level division-of-labor framework.