arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26926cs.CLcs.AIcs.HCcs.LG

专家在LLM分歧处崛起:利用跨模型分歧定位专家精力于大规模标注的LLM代码本修订

Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation

发表机构宾夕法尼亚州立大学 · 康奈尔大学
查看机构详情
  • The Pennsylvania State University(宾夕法尼亚州立大学)
  • Cornell University(康奈尔大学)

机构由 AI 辅助整理,请以论文原文为准。

Zeyu He, Zhuqian Zhou, Kirk Vanacore, Rene F. Kizilcec, Ting-Hao 'Kenneth' Huang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究利用LLM分歧定位专家精力,通过理由标注等反馈方式修订代码本,在数千个辅导会话中达到64.9%的标注准确率,将数月修订缩短至数天。

中文摘要 AI 辅助

大规模文本标注通过AI标注者遵循的代码本,将专家见解应用于数百万份文档。然而,开发一个健壮的代码本需要数月时间。大型语言模型(LLMs)可以通过将早期代码本应用于数据、突出显示LLM强烈分歧的案例,并征求专家反馈来解决这些问题,从而加速这一过程。我们研究了专家为LLM代码本修订提供反馈的三种方式:(i)编辑由跨LLM分歧驱动的LLM生成修订(代码本验证),(ii)回答关于LLM分歧的问题(问答),以及(iii)用理由标注分歧案例(理由标注)。在数千个辅导会话记录上的实验表明,理由标注在专家标签上的LLM标注准确率最高(64.9%),优于专家修订的代码本(57.8%)。最佳问答设置也优于该代码本(60.5%)。我们的工作表明,LLMs可用于战略性地定位专家注意力,将数月的代码本修订缩短至数天,且不牺牲标注性能。

英文摘要

Large-scale text annotation brings expert insight to millions of documents, often through a codebook that AI annotators follow. Developing a robust codebook, however, takes months. Large language models (LLMs) could speed this process by applying an early codebook to the data, surfacing cases with strong LLM disagreement, and eliciting expert feedback to address them. We examined three ways experts can provide feedback for LLM codebook revision: (i) editing LLM-generated revisions driven by cross-LLM disagreement (Codebook Verifying), (ii) answering questions about LLM disagreements (Question Answering), and (iii) labeling disagreement cases with rationales (Rationale Labeling). Experiments on thousands of tutoring-session transcripts show that Rationale Labeling yielded the highest LLM-labeling accuracy (64.9%) against expert labels, outperforming the expert-revised codebook (57.8%). The best Question Answering setting also outperformed it (60.5%). Our work shows that LLMs can be used to strategically target expert attention, shortening months of codebook revision to days without sacrificing labeling performance.

↑