arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

边界一致,标签不一致:标注者内部动态作为一种训练数据

Boundaries Agree, Labels Do Not: Intra-Annotator Dynamics as a Kind of Training Data

Marharyta Shvets

arXiv 2610.04370首次发表:更新:

发表机构

Eötvös Loránd University (ELTE)(厄特沃什·罗兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过比较人类与LLM在多次标注苏美尔神话时的分段与标签稳定性,提出标注者内部动态可作为数据质量与污染检测指标。

AI 中文摘要

数据质量如今在训练语言模型方面与计算能力同等重要。大量训练数据来自对文本的人工标注,而解释性标注没有能够确定何为“准确”的基准真相。针对这一问题,有两类研究工作。一类将标注者合并为“基准真相”,并衡量他们彼此之间的一致性;另一类则将他们的分歧视为一种信号。这两类方法都是在同一时间点对不同的人进行比较。我们衡量的是另一件事:一位读者在多次阅读同一文本时,其自身阅读结果的可复现程度。一位专家人类读者和三个LLM家族对三部苏美尔神话进行了分段,并用七种状态之一标注了每个段落的因果功能。在相隔数月的多次运行中,人类在每部神话中都在大致相同的位置切割文本,但对段落的命名却不同。模型没有表现出这种一致的模式:它们在两个层面之间的差距在某些神话中为正,在另一些神话中为负,且其大小有所变化。人类的标签变化并非随机:多次运行经历了大致相同的功能,但起始位置相差一步,而模型运行则在相同位置开始这些功能。我们认为,这种模式是一种可用的数据质量度量标准和污染检查手段:一种标签与其边界一样稳定、且其功能同步开始的“人类”标注,看起来就像模型的标注。

英文摘要

Data quality now matters as much as compute for training language models. Much training data comes from human annotation of text, and interpretive annotation has no ground truth that could settle what is "accurate". Two lines of work respond to this. One combines annotators into a "ground truth" and measures how well they agree with each other; the other treats their disagreement as a signal. Both compare different people at one point in time. We measure something else: how well one reader reproduces their own reading of the same text over time. One expert human reader and three LLM families segmented three Sumerian myths and labelled the causal function of each segment with one of seven states. Across runs months apart, the human cut the text in much the same places but named the segments differently, in every myth. The models show no such consistent pattern: their gap between the two layers is positive in some myths and negative in others, and its size varies. The human's label changes are not random: the runs go through much the same functions but start them one step apart, while model runs start them at the same places. We argue that this pattern is a usable measure of data quality and a contamination check: a "human" annotation whose labels are as stable as its boundaries, and whose functions start in sync, looks like a model's.

Comments8 pages, 1 figure, 5 tables. Code and data: https://github.com/Malificenta883/intra-annotator-dynamics

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑