发表机构
RAIC Labs(RAIC实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对单类发现问题,提出FunnelAL主动学习系统,将工业推荐系统的多阶段漏斗架构用于数据标注,分解问题为级联阶段,经多轮迭代完善,在图像分类基准测试中表现优异,为多阶段推荐与主动学习搭建桥梁。
AI 中文摘要
我们提出了FunnelAL,一种用于单类发现的先检索后排序主动学习系统,它将工业推荐系统的多阶段漏斗架构应用于数据标注。大规模监督学习面临两个挑战:在海量语料库中高效找到相关样本,以及当嵌入不能清晰分离类别时从视觉上易混淆的负样本中区分真阳性。传统主动学习虽提供了降低标注成本的框架,但将样本选择视为单阶段过程,无法有效应对这两个挑战。FunnelAL将问题分解为级联阶段。从单个正例和负例开始,系统依次进行:基于嵌入的检索评分,将语料库缩小到可管理的候选集;精度触发的排序阶段,在批量精度高时利用学习到的排序器(RankNet),一旦回报减少则自动融入基于委员会的探索(QBC);以及来自注释者标签的反馈,在后续迭代中完善两个阶段。我们在三个不同的图像分类基准上进行评估。在有完美注释者的情况下,FunnelAL在所有三个基准上都获得了最佳的最终F1、最佳的标注效率(在AULC中排名第一)和最少的标注轮次。最新的单类发现方法(GAL、PF-MA)充其量只能匹配其最终质量,且标注成本始终更高。在实际比例的注释者标注错误情况下,FunnelAL仍排名第一或在统计上并列第一,而基于经典不确定性的方法退化速度快两到三倍。我们的工作在多阶段推荐系统和主动学习之间架起了一座具体的桥梁。
英文摘要
We present FunnelAL, a retrieve-then-rank active learning system for single-class discovery, which adapts the multi-stage funnel architecture of industrial recommender systems to data annotation. Large-scale supervised learning faces two challenges: efficiently finding relevant samples in a massive corpus, and distinguishing true positives from visually confusable negatives when embeddings do not cleanly separate classes. Conventional active learning offers a principled framework for reducing annotation cost, yet it treats sample selection as a single-stage process that addresses neither challenge efficiently. FunnelAL decomposes the problem into cascaded stages. Starting from a single positive and negative example, the system iterates through: (1) embedding-based retrieval scoring that narrows the corpus to a manageable candidate set; (2) a precision-triggered ranking stage that exploits a learned ranker (RankNet) while batch precision remains high, then automatically blends in committee-based exploration (QBC) once returns diminish; and (3) feedback from the annotator's labels that refines both stages in subsequent iterations. We evaluate on three diverse image classification benchmarks. With a perfect annotator, FunnelAL attains the best final F1 on all three benchmarks, the best annotation efficiency (first in AULC), and the fewest annotation rounds. The most recent single-class discovery methods (GAL, PF-MA) at best match its final quality, and only at consistently higher labeling cost. Under annotator labeling errors at realistic rates, FunnelAL remains first or statistically tied for first while classical uncertainty-based methods degrade two to three times faster. Our work provides a concrete bridge between multi-stage recommender systems and active learning.
Comments15 pages, 6 figures, 3 tables