arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可靠但对设计敏感:LLM标注中的工具不确定性

Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation

Thomas Reiter, Christoph Kern, Fedor Miasnikov, Sofiia Nikolenko, Rob Chew, Stephanie Eckman, Frauke Kreuter

arXiv 2609.35824首次发表:更新:

发表机构

Amazon; Ludwig Maximilian University of Munich; RTI International; University of Maryland(亚马逊; 慕尼黑大学; RTI国际; 马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示LLM标注在不同任务设计下标签不稳定,提出工具不确定性概念,并证明仅靠重复设置或置信度分数无法替代跨设计比较来评估可靠性。

AI 中文摘要

大型语言模型(LLMs)在一种设置下能给出可靠的标签,但当研究人员做出其他合理的设计选择时,这些标签可能会改变。我们测试了7个LLM、12种任务设计、3次独立运行,以及3000条被标记为攻击性语言和仇恨言论的推文。重复相同的模型和任务设计产生了高度一致性(中位Fleiss' κ = 0.91)。当我们改变相同推文的任务设计时,一致性下降(中位Cohen's κ = 0.76)。与仅采样方差相比,任务设计和模型选择使估计流行率的方差增加了攻击性语言76.7倍和仇恨言论110.6倍。LLM任务设计间的变异达到560-572个基点,而五个人类工具版本间的变异为270-331个基点。置信度分数并未解决此问题。它们与重复模型输出的跟踪比对人类标签的一致性更紧密,且将六条推文分组在一个提示中将平均攻击性语言置信度降低了660个基点。我们将由任务设计和模型选择引起的变异称为工具不确定性。研究人员只能通过比较合理的任务设计来测量它。重复一种设置或依赖置信度分数不能替代该测试。

英文摘要

Large language models (LLMs) can give reliable labels under one setup yet change those labels when researchers make other reasonable design choices. We tested seven LLMs, 12 task designs, three independent runs, and 3,000 tweets labeled for offensive language and hate speech. Repeating the same model and task design produced high agreement (median Fleiss' $κ= 0.91$). Agreement fell when we changed the task design for the same tweets (median Cohen's $κ= 0.76$). Task design and model choice increased the variance of estimated prevalence by factors of 76.7 for offensive language and 110.6 for hate speech compared with sampling variance alone. Variation across LLM task designs reached 560-572 basis points, compared with 270-331 basis points across five human instrument versions. Confidence scores did not solve this problem. They tracked repeated model outputs more closely than agreement with human labels, and grouping six tweets in one prompt lowered mean offensive-language confidence by 660 basis points. We call the variation caused by task design and model choice instrument uncertainty. Researchers can measure it only by comparing reasonable task designs. Repeating one setup or relying on confidence scores cannot replace that test.

CommentsAccepted to "3rd Workshop on Uncertainty-Aware NLP" @ EMNLP 2026 (archival)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑