arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向K-12教育工作者AI使用的人机协同归纳编码方法

Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use

Alex Liu, Min Sun, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He

arXiv 2607.28889首次发表:更新:

发表机构

College of Education, University of Washington; Colleague AI(华盛顿大学教育学院; Colleague AI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出多阶段人机协同编码流程,将LLM作为标注工具辅助K-12教育工作者交互语料库编码,保留人类概念权威,构建含72个条目的分层编码本,为定性研究提供可审计模板。

AI 中文摘要

定性研究人员日益面临交互语料库规模超出手动编码处理能力的情况,因此常提出将大型语言模型(LLM)作为分析助手。关键问题并非LLM能否参与定性分析,而是其参与程度、适用阶段及所需保障措施。本文详细阐述了多阶段人机协同流程,该流程采用开放编码、主轴编码和选择性编码方法,从K-12教育工作者与生成式AI平台交换的45000条消息中构建分层编码本。在三个阶段中,LLM大规模生成候选标签和结构化注释,而人类研究人员保留对类别定义、合并决策和解释框架的概念权威。随后通过系统人类编码对该工具进行测试:三名具备教育领域专业知识的训练有素编码员将编码本应用于2560条独立样本消息,采用适用于多标签注释的集值一致性测量方法通过迭代校准建立可靠性,并补充了LLM辅助阶段未发现的5个编码。最终编码本包含6个领域下19个类别中的72个条目。本文反思了该流程所需的方法决策,包括选择会话分析单位、将LLM视为标注工具而非解释主体、多标签编码下的编码员间一致性测量,以及人类领域专业知识保持决定性的条件。该流程为考虑在编码本开发中使用LLM辅助同时保留人类解释权威的定性研究人员提供了可审计模板。

英文摘要

Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models (LLMs) are frequently proposed as analytic assistants. The open questions are not whether LLMs can participate in qualitative analysis but to what extent, in what phases, and under what safeguards. This article provides a detailed procedural account of a multi-phase human-LLM collaborative pipeline that adapted open, axial, and selective coding to develop a hierarchical codebook from 45,000 messages exchanged between K-12 educators and a generative AI platform. Across three phases, LLMs generated candidate labels and structured annotations at scale, while human researchers retained conceptual authority over category definitions, merging decisions, and interpretive frameworks. The resulting instrument was then tested through systematic human coding, in which three trained coders with educational domain expertise applied the codebook to an independent sample of 2,560 messages, established reliability through iterative calibration using set-valued agreement measures appropriate for multi-label annotation, and extended the instrument with five codes that the LLM-assisted phases had not surfaced. The final codebook comprises 72 items within 19 categories and six domains. We reflect on the methodological decisions the pipeline required, including the choice of a conversational unit of analysis, the treatment of the LLM as a labeling instrument rather than an interpretive agent, the measurement of intercoder agreement under multi-label coding, and the conditions under which human domain expertise remained decisive. The account is offered as an auditable template for qualitative researchers considering LLM assistance in codebook development while preserving human interpretive authority.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑