发表机构
Amazon; Sapienza University of Rome; Université Côte d’Azur; CNRS; Inria; I3S(亚马逊; 罗马萨皮恩扎大学; 蔚蓝海岸大学; 法国国家科学研究中心; 法国国家信息与自动化研究所; 信息、信号与智能系统实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出预算感知路由框架QUORUM,通过特征信号估计实例难度、结合多标注者结果,在固定预算下动态分配任务,提升标注质量并降低成本。
AI 中文摘要
数据标注仍是自然语言处理的核心瓶颈,需要人力来大规模获取高质量标签。虽然大型语言模型(LLMs)提供了快速且具成本效益的替代方案,但其可靠性高度依赖实例:它们在简单输入上表现良好,但在需要细致推理或上下文理解的示例上常常失效。本研究针对这一挑战,提出了QUORUM(基于多个标注者的质量优化路由,QUality-Optimized Routing Using Multiple annotators),这是一种预算感知的路由框架,在固定标注预算下动态将每个实例分配给人类或LLM标注者。与依赖模型置信度或不确定性估计的现有方法不同,QUORUM利用基于特征的信号来估计实例难度,并支持每个实例的多个标注,通过基于一致性的奖励将它们结合以提高可靠性。我们在英语和多语言环境中的各类封闭式和开放式标注任务上对QUORUM进行评估,结果显示,与竞争方法相比,QUORUM可将标注质量提升高达34.4%,同时降低8.8%的成本。代码可在该https链接获取。
英文摘要
Data annotation remains a central bottleneck in natural language processing, requiring human effort to obtain high-quality labels at scale. While Large Language Models (LLMs) offer a fast and cost-effective alternative, their reliability is highly instance-dependent: they perform well on simple inputs but often fail on examples requiring nuanced reasoning or contextual understanding. In this work, we address this challenge with QUORUM (QUality-Optimized Routing Using Multiple annotators), a budget-aware routing framework that dynamically assigns each instance to human or LLM annotators under a fixed annotation budget. Unlike prior approaches relying on model confidence or uncertainty estimates, QUORUM leverages feature-based signals to estimate instance difficulty and supports multiple annotations per instance, combining them through agreement-based rewards to improve reliability. We evaluate QUORUM across diverse closed- and open-ended annotation tasks in English and multilingual settings, and QUORUM improves annotation quality by up to 34.4% while reducing costs by 8.8% over competing methods. Code can be found at https://github.com/amazon-science/QUORUM.
Comments4 figures, 18 pages