发表机构
The University of Melbourne; Jimei University(墨尔本大学; 集美大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对目标检测主动学习,提出基于类原型距离的评分信号,仅需一次前向传播,以少量参数提升标注选择效率,在多个数据集上优于后验概率并匹敌集成方法。
AI 中文摘要
在新环境中部署深度目标检测器,其限制主要不在于架构,而在于标注该环境数据的成本。主动学习通过选择要标注的图像来降低成本,而选择的好坏取决于用于给未标注图像打分的信号。该信号通常是类别后验概率,它计算廉价但校准不佳;或者是多个模型或多个随机前向传播之间的分歧,后者效果更好,但会使得推理次数在远大于标注集的池上成倍增加。我们提出一种比后验概率更丰富、但仍可通过单个网络的一次前向传播读取的信号。在训练目标中加入一个有监督对比项,以塑造一个逐对象嵌入空间,在该空间中距离编码类别成员关系;一个未标注检测的得分取决于其预测类别所占据区域的距离,并以置信度加权。该准则不需要集成、不需要辅助预测器、也不需要重复推理,其全部成本为289万个参数,比裸检测器增加8.3%。在PASCAL VOC和MS-COCO上,它在每一轮选择中都优于同一检测器的后验概率,最多高出1.08%的mAP50,而运行间偏差为0.02%至0.18%;并且它与集成和蒙特卡洛dropout准则保持竞争力,后者每张未标注图像需要三到五十次前向传播。实验使用单阶段检测器,所比较的准则均在该检测器上报告其结果,从而将选择决策与检测器的强度分离开来。
英文摘要
Deploying a deep object detector in a new setting is limited less by architecture than by the cost of annotating data from that setting. Active learning lowers the cost by choosing which images to label, and the choice is only as good as the signal used to score an unlabeled image. That signal is usually the class posterior, which is cheap but poorly calibrated, or the disagreement across several models or several stochastic passes, which is better but multiplies inference over a pool far larger than the labeled set. We propose a signal richer than the posterior yet still read from one forward pass of one network. A supervised contrastive term added to the training objective shapes a per-object embedding space in which distance encodes class membership, and an unlabeled detection is scored by how far it lies from the region occupied by its predicted category, weighted by its confidence. The criterion needs no ensemble, no auxiliary predictor and no repeated inference, and its entire cost is 2.89M parameters, an increase of 8.3% over a bare detector. On PASCAL VOC and MS-COCO it beats the posterior of the same detector in every round in which a selection is made, by up to 1.08% mAP50 against run to run deviations of 0.02% to 0.18%, and it stays competitive with ensemble and Monte Carlo dropout criteria costing three to fifty forward passes per unlabeled image. Experiments use the single-stage detector under which the compared criteria report their results, so that the selection decision is isolated from the strength of the detector.