arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22133cs.CL

LLM与人类标注的观测等价性

Observational Equivalence of LLM and Human Annotation

Kentaro Nakamura, Jing Ling Tan, George Yean

首次发表
浏览论文内容

中文总结 AI 辅助

本文证明LLM与人类在文本标注质量上观测等价,分歧源于文本模糊性,建议通过LLM分歧识别困难案例并完善编码规则,以发挥速度和成本优势。

中文摘要 AI 辅助

在本文中,我们证明在标注质量方面,LLM与人类编码是观测等价的:最近的LLM与专家编码者的一致性程度与专家之间的一致性程度相当。我们通过对来自14项同行评审政治学研究的文本分类任务进行复制来证明这一点,在这些任务中,十个LLM、三位人类专家和165名众包工作者使用相同的编码手册独立对相同文本进行分类。我们发现,这种等价性是由文本和编码规则中的模糊性驱动的。当LLM与专家意见不一致时,专家之间也更可能意见不一致,而澄清编码规则可以减少专家和足够有能力的LLM之间的分歧。因此,仅基于标注质量而偏好人类编码几乎没有经验依据,而LLM在速度和成本方面具有显著优势。因此,我们认为文本标注的核心挑战不再是选择人类还是机器编码者,而是制定能最小化模糊性的编码规则,并处理剩余的模糊性。为此,我们建议利用LLM之间的分歧来识别困难案例并完善编码手册,并且当无法为每个文本定义唯一标注时,我们为下游推断开发了具有模糊性感知的界限。

英文摘要

In this paper, we show that LLM and human coding are observationally equivalent in terms of annotation quality: recent LLMs agree with expert coders at rates comparable to those observed among experts themselves. We demonstrate this through replications of text-classification tasks from 14 peer-reviewed political science studies, in which ten LLMs, three human experts, and 165 crowdsourced workers independently classify the same texts using identical codebooks. We find that this equivalence is driven by ambiguity in the texts and coding rules. When LLMs disagree with experts, experts are also more likely to disagree with one another, and clarifying coding rules reduces disagreement among both experts and sufficiently capable LLMs. Thus, there is little empirical basis for preferring human coding on the basis of annotation quality alone, while LLMs offer substantial advantages in speed and cost. We therefore argue that the central challenge of text annotation is no longer choosing between human and machine coders, but developing coding rules that minimize ambiguity and accounting for the ambiguity that remains. To this end, we propose using disagreement across LLMs to identify difficult cases and refine codebooks, and we develop ambiguity-aware bounds for downstream inference when a unique annotation cannot be defined for every text.

发表机构

  • Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

↑