数据标注即测量
Data Annotation as Measurement
浏览论文内容
中文总结 AI 辅助
本文将数据标注重新定义为测量,通过文献综述和访谈开发标注问题诊断框架,明确标注问题来源并给出超越一致性的质量评估方法,为提升AI标注数据质量提供概念基础。
中文摘要 AI 辅助
现代AI系统依赖标注数据,但标注很少被视为其本质上的测量行为。相反,标注质量通常被简化为一致性:如果多个标注者对某一数据实例赋予相同标注,则认为该标注质量高。然而,一致性并不能证明标注是否有效捕获了它们旨在代表的潜在概念。在本文中,我们主张应将数据标注理解为一个测量问题。与其他形式的测量类似,标注需要定义一个概念、通过工具将其操作化、应用该工具,并评估所得测量结果的可靠性与有效性。基于对标注质量研究的文献综述(样本量N=132)以及对标注团队成员的半结构化访谈(样本量N=10),我们开发了一套用于诊断和纠正标注问题的框架。首先,我们梳理了标注流程中的关键决策点,包括任务设计、标注者管理、质量评估、质量改进与裁决,这些决策点会影响标注结果。其次,我们确定了五类不同的标注问题来源:错误、歧义、不可行、主观性以及标注者身份。在结果层面表现相似的标注问题,根据其来源往往需要不同的流程层面干预措施。最后,我们将测量理论转化为标注团队的实用指南,展示了如何超越仅依赖一致性的方式来评估可靠性与有效性。通过将标注重新定义为测量,我们为提高AI研究与实践中所用标注数据的质量提供了概念基础。
英文摘要
Modern AI systems depend on annotated data, but annotation is rarely treated as the act of measurement that it is. Instead, annotation quality is commonly reduced to agreement: if multiple annotators assign the same annotation to a data instance, the annotations are taken to be high-quality. Yet agreement does not establish whether annotations validly capture the underlying concept they are meant to represent. In this paper, we argue that data annotation should be understood as a measurement problem. Like other forms of measurement, annotation requires defining a concept, operationalizing it through an instrument, applying that instrument, and evaluating the reliability and validity of the resulting measurements. Drawing on a literature review of annotation quality research (N=132) and semi-structured interviews with annotation team members (N=10), we develop a framework for diagnosing and correcting annotation issues. First, we map key decision points across annotation processes - including task design, annotator management, quality assessment, quality improvement, and adjudication - that shape annotation outcomes. Second, we identify five distinct sources of annotation issues: error, ambiguity, impossibility, subjectivity, and annotator identity. Annotation problems that appear similar at the level of outcomes often require different process-level interventions based on their sources. Finally, we translate measurement theory into practical guidance for annotation teams, showing how assessments of reliability and validity can move beyond agreement alone. By reframing annotation as measurement, we offer a conceptual foundation for improving the quality of annotated data used in AI research and practice.