AI编码者如何讨论、分歧与达成共识:基于LLM的定性编码面临的挑战与机遇
How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding
浏览论文内容
中文总结 AI 辅助
本研究量化多智能体LLM在定性编码中的有效性,发现编码手册长度、数据相似性和智能体分歧影响准确性,激烈未解决分歧反而提高准确率,并提出设计建议与开源框架。
中文摘要 AI 辅助
人工智能在多编码者定性编码中的效用已被广泛讨论,但关于其在何种情境下能够可靠运行的实证证据仍然很少。我们通过量化多智能体LLM编码在不同定性数据集上的有效性来填补这一空白,揭示了调节编码结果的关键情境因素和结构因素。我们开发了一个基于文献的基线流程,使AI智能体能够独立编码、辩论并调和分歧。结果表明,编码准确性取决于编码手册长度、定性数据相似性和智能体分歧等因素。值得注意的是,智能体之间激烈且未解决的分歧导致了更高的准确性。我们的分析表明,虽然LLM模拟了许多人类讨论行为,但它们缺乏对情境的适应性响应。基于这些发现,我们为构建自动化编码系统提供了设计建议。我们开源的人工智能讨论数据集和方法论框架为推进人工智能介导的自动化主题分析设计奠定了基础。
英文摘要
The utility of AI in multi-coder qualitative coding has been widely discussed, yet little empirical evidence exists to delineate the contexts in which it performs reliably. We address this gap by quantifying the effectiveness of multi-agent LLM coding across varied qualitative datasets, revealing key contextual and structural factors that mediate coding outcomes. We developed a literature-informed baseline pipeline that enables AI agents to independently code, debate, and reconcile disagreements. Results revealed that coding accuracy depends on factors such as codebook length, qualitative data similarity, and agent disagreement. Notably, intense and unresolved debates between agents led to higher accuracy. Our analysis showed that while LLMs emulate many human discussion behaviors, they lack adaptive responsiveness to context. From these findings, we offer design recommendations for building automated coding systems. Our open-source AI discussion dataset and methodological framework lay the groundwork for advancing the design of AI-mediated automated thematic analysis.
发表机构
- Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。