长尾分布中少即是多:面向标注高效密集预测的阶段自适应样本选择
Less Is More in the Long Tail: Stage-Adaptive Sample Selection for Annotation-Efficient Dense Prediction
浏览论文内容
中文总结 AI 辅助
针对长尾密集预测中标注成本高的问题,提出阶段自适应样本选择框架SASS,结合自监督梯度评分、类别重平衡与阶段自适应获取,以40%标注预算恢复98.3%性能,并展现少即是多模式。
中文摘要 AI 辅助
深度学习性能通常随着训练数据的增加而提升,然而在大规模密集预测任务中,由于类别分布呈长尾特性,像素级或体素级标注成本极高,这种扩展从根本上受到标注成本的制约。我们提出SASS(阶段自适应样本选择),一种用于长尾密集预测中基于池的主动学习的阶段自适应数据选择框架。SASS结合了三个组件:无标签的自监督梯度评分、带有验证驱动反馈的先验引导类别重平衡,以及与模型训练动态对齐的阶段自适应获取。这种设计在梯度评分过程中避免了候选真实掩码,同时使获取过程对长尾不平衡和不断演化的表示具有响应性。我们在一个包含超过100,000个样本、涵盖108个解剖结构的多模态3D医学分割测试平台上评估了SASS。SASS以40%的训练池标注预算恢复了全数据集性能的98.3%,比BADGE高出5.1个百分点。此外,SASS表现出统计学上支持的少即是多模式,在Hard组级别以及结构级别上(针对胰腺和胆囊)超越了全数据集训练。更广泛地说,SASS表明标注高效学习不仅取决于选择哪些样本,还取决于标注预算如何在类别间分配,以及模型派生分数何时开始指导选择。
英文摘要
Deep learning performance generally improves with increasing training data, yet this scaling is fundamentally constrained by annotation cost in large-scale dense prediction tasks with long-tailed category distributions, where pixel- or voxel-level annotation is prohibitively expensive. We propose SASS (Stage-Adaptive Sample Selection), a stage-adaptive data-selection framework for pool-based active learning in long-tailed dense prediction. SASS combines three components: label-free self-supervised gradient scoring, prior-guided category rebalancing with validation-driven feedback, and stage-adaptive acquisition aligned with model training dynamics. This design avoids candidate ground-truth masks during gradient scoring while making acquisition responsive to long-tail imbalance and evolving representations. We evaluate SASS on a multimodal 3D medical segmentation testbed comprising over 100,000 samples spanning 108 anatomical structures. SASS recovers 98.3% of full-dataset performance with a 40% training-pool annotation budget, outperforming BADGE by 5.1 percentage points. Moreover, SASS exhibits a statistically supported less-is-more pattern, surpassing full-dataset training at the Hard-group level and, at the structure level, for the pancreas and gallbladder. More broadly, SASS shows that annotation-efficient learning depends not only on which samples are selected, but also on how the annotation budget is distributed across categories and when model-derived scores begin to guide selection.
发表机构
- Fudan University(复旦大学)
- University of Exeter(埃克塞特大学)
机构由 AI 辅助整理,请以论文原文为准。