AI 中文总结
研究多类标签聚合问题,提出AHEAD框架,通过图神经网络学习交叉注释器上下文,聚合特征得到注释器嵌入并解码为混淆矩阵,结合高置信度注释器缓解训练问题,实验表明该方法显著提高标签准确性和可扩展性。
AI 中文摘要
众包标注为自然语言处理、计算机视觉和视频等领域提供了有价值的标注数据。标签聚合旨在从噪声和有偏差的注释中推断潜在的真实标签,关键在于注释器可靠性估计。现有方法存在瓶颈,多数个体注释器仅标注一小部分任务,准确估计注释器很棘手。本文聚焦更具挑战性的多类标签聚合,提出AHEAD框架,通过利用总体级数据推进注释器可靠性估计。具体而言,AHEAD先通过图神经网络学习高维交叉注释器上下文,聚合个体级注释器特征与上下文信息得到多视图、互补的注释器嵌入,再解码为可解释的特定注释器混淆矩阵以拟合观察到的标签。还制定了包含高置信度注释器的复合目标以缓解先前模型面临的无监督训练问题。在10个跨NLP、CV、视频和音频的真实世界数据集上的实验表明,AHEAD显著提高了标签准确性,平均准确率从68.75%提高到73.23%,最佳情况下增益高达14.9%。同时,在最大数据集上的可扩展性实验进一步证明了该方法的总体优越性。
英文摘要
Crowdsourced labeling provides valuable labeled data for domains across natural language processing, computer vision, and video. Label aggregation aims to infer latent true labels from noisy and biased annotations, with the key lying in annotator reliability estimation. Despite promising progress, existing approaches struggle with one real-world bottleneck: most individual annotators label only a small subset of tasks, making accurate annotator estimation highly intractable. In this paper, we focus on the considerably more challenging multi-class label aggregation and propose AHEAD (cross-Annotator learning and High-confidEnce Annotator-guideD label aggregation), a cross-annotator learning framework that advances annotator reliability estimation by leveraging the population-level data. Specifically, AHEAD first learns high-dimensional cross-annotator contexts via a graph neural network, deriving multi-view, complementary annotator embeddings by aggregating individual-level annotator features with contextual information. These embeddings are then decoded into interpretable annotator-specific confusion matrices to fit the observed labels. We formulate a composite objective incorporating high-confidence annotators to alleviate the unsupervised training issues faced by prior models. Experiments on 10 real-world datasets spanning NLP, CV, Video, and Audio show that AHEAD substantially improves label accuracy, increasing average accuracy from 68.75% to 73.23%, with gains of up to 14.9% in the best case. Meanwhile, scalability experiments on the largest dataset further demonstrate the overall superiority of our method.