AI 中文总结
该研究针对冷启动主动学习方法适配性不足的问题,提出基于最优传输的ε-AS算法,在六个数据集上实现最优性能,在ImageNet-1k上较ActiveFT提升准确率且缩短选择时间
AI 中文摘要
冷启动主动学习(CSAL)旨在从未标记池中选择有价值的子集,且无需任何先验知识或人工协助。现有方法基于典型性、覆盖度或多样性采取不同路径,每种方法都有自身的归纳偏置,因此在部分任务上表现良好,而在其他任务上表现较差。我们认为真正的挑战并非设计另一种选择启发式,而是让CSAL自动适配手头的数据和任务。为此,我们从最优传输视角重新审视CSAL:首先,提出一种广义传输选择框架,揭示现有方法的共享分配结构,且可精确涵盖代表性公式;其次,引入理论分析,刻画由熵正则化控制的权衡关系,并建立冷启动选择的任务不可知极小极大界,这些结果为将正则化强度适配到未标记数据提供了原则性基础;最后,推导数据自适应正则化规则,提出一种基于Sinkhorn的新型CSAL算法,称为ε自适应选择(ε-AS)。在六个公开数据集和多个标注预算上开展的大量实验表明,ε-AS始终达到最优性能;在ImageNet-1k上,它相比ActiveFT将平均准确率提升1.29%,同时将选择时间减少56.2%。代码将在此处httpsURL发布
英文摘要
Cold-Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance. Existing methods take diverse routes based on typicality, coverage, or diversity. Each rests on its own inductive bias and therefore performs well on some tasks yet poorly on others. We argue that the real challenge is not to design yet another selection heuristic, but to make CSAL adapt automatically to the data and task at hand. To this end, we revisit CSAL through the lens of optimal transport. First, we propose a generalized transport selection framework that reveals the shared allocation structure of existing methods and exactly subsumes representative formulations. Second, we introduce a theoretical analysis that characterizes the trade-off controlled by entropic regularization and establishes a task-agnostic minimax bound for cold-start selection. These results provide a principled foundation for adapting the regularization strength to the unlabeled data. Third, we derive a data-adaptive regularization rule and present a novel Sinkhorn-based CSAL algorithm, termed $ε$-Adaptive Selection ($ε$-AS). Extensive experiments on six public datasets and multiple annotation budgets show that $ε$-AS consistently achieves state-of-the-art performance. On ImageNet-1k, it improves the average accuracy over ActiveFT by 1.29% while reducing selection time by 56.2%. Code will be released at https://github.com/Z-yiwei/OT-CSAL