arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24622stat.MLcs.LG

一种不平衡标签聚合模型:聚焦少数类检测

A Model for Imbalanced Label Aggregation: A Focus on Minority-Class Detection

Gabriel Singer, Samuel Gruffaz, Olivier Vo Van, Nicolas Vayatis, Argyris Kalogeratos

首次发表
浏览论文内容

中文总结 AI 辅助

研究不平衡众包中依赖类别的注释器准确性问题,引入结合项目难度与注释器能力的生成聚合模型,在33个真实世界众包数据集上评估,该模型在少数类召回率上表现出色,平衡准确率也具竞争力。

中文摘要 AI 辅助

我们研究不平衡众包,重点关注依赖类别的注释器准确性。据我们所知,尽管在实际检查系统中,最重要的标签也是最罕见的标签,但这种情况仍未得到充分探索。在这种情况下,注释器可能在两类上都可靠、都不可靠、是多数类专家或少数类专家。现有模型仅部分解决此问题。为填补众包中不平衡数据集的这一空白,我们引入一种将项目难度与依赖类别的注释器能力相结合的生成聚合模型。该模型允许注释器能力和项目难度因类而异。我们还在类不平衡设置中重新审视了孔多塞陪审团定理。我们在33个真实世界众包数据集上评估了我们的模型,涵盖图像和文本等多类任务以及两种大规模情况。在这些不同设置下,我们的模型始终实现最高的少数类召回率,同时在平衡准确率方面具有竞争力。

英文摘要

We study imbalanced crowdsourcing with a focus on class-dependent annotator accuracy, a setting that, to the best of our knowledge, remains relatively underexplored despite its importance in real-world inspection systems where the labels of greatest operational importance are also the rarest ones. In this setting, annotators may be reliable on both classes, unreliable on both classes, majority-class specialists, or minority-class specialists. Existing models only partially address this problem: they either capture class-dependent errors but ignore item difficulty, or they model item difficulty without capturing class-dependent errors. To fill this gap for imbalanced datasets in crowdsourcing, we introduce a generative aggregation model combining item difficulty with class-dependent annotator competence. The model allows both annotator abilities and item difficulties to vary across classes. We then revisit Condorcet's Jury Theorem in the class-imbalanced setting. We also show that majority voting asymptotically preserves the underlying class proportion. We evaluate our model on $33$ real-world crowdsourcing datasets, covering multiclass tasks such as images and text, as well as two large-scale regimes: large-scale annotation datasets, with many annotations per item, and large-scale item datasets, with a large number of annotated instances. Across these diverse settings, our model consistently achieves the highest minority recall while remaining competitive in balanced accuracy, making it particularly relevant when rare-label recovery is the primary objective.

发表机构

  • SNCF(法国国家铁路公司)
  • Tampere University(坦佩雷大学)

机构由 AI 辅助整理,请以论文原文为准。

↑