批量模式主动学习的软策略选择
Soft Strategy Selection for Batch-Mode Active Learning
浏览论文内容
中文总结 AI 辅助
针对主动学习中预先选择单一采集策略的难题,提出FractAL方法,在单批次内通过影响函数归因和在线镜像下降实现多策略软选择,在7种设置上均匹配或优于基线,为高通量实验提供可靠策略选择。
中文摘要 AI 辅助
主动学习在实际部署中通常迫使从业者在任何数据被标注之前就选择一种采集策略。这是一项艰巨的任务:策略性能在不同设置(如数据集、替代模型)之间差异很大,且无法在不部署的情况下进行评估。现有的策略选择方法在每轮中从策略组合中探索一种策略,并通过赌博机反馈或模型重训练来识别最优策略。因此,许多采集轮次被用于探索策略,而非收集最具信息量的数据。这种开销是主动学习驱动的高通量实验(如基因扰动筛选和定向进化)设计的主要障碍,在这些实验中,主动学习运行仅由少数几轮组成,且批次规模很大。这种机制允许一种自然的替代方案:在单个批次内使用多种主动学习策略来采集数据。我们将此称为软策略选择,并引入FractAL,一种专门为此任务设计的方法。FractAL使用基于影响函数的数据归因来推断每种策略的奖励,无需额外重训练,然后使用在线镜像下降计算策略组合中每种策略的预算份额。我们在涵盖分类、回归和基因扰动效应预测的7种设置上对FractAL进行了基准测试。结果强调策略选择是一个难题:每种现有方法在至少一种设置上的表现都差于随机采样。然而,FractAL在所有7种设置上均匹配或优于所有基线,包括随机采样。其分配将预算集中在组合中最强的策略上,同时削减最弱的策略。因此,对于实际部署(其中最优策略事先未知),FractAL是一个可靠的选择,这是使主动学习适用于高通量实验和现代科学发现的重要一步。
英文摘要
Real-world deployment of active learning typically forces practitioners to choose an acquisition strategy before any data is labeled. This is a daunting task: strategy performance varies widely across settings (e.g. datasets, surrogate models) and cannot be assessed without deployment. Existing strategy selection methods explore one strategy from a portfolio at each round and identify the optimal one using bandit feedback or model retraining. Many acquisition rounds are therefore spent exploring strategies rather than collecting the most informative data. Such overhead is a major barrier to AL-driven design of high-throughput experiments, such as genetic perturbation screens and directed evolution, where AL runs consist of only a few rounds with large batch sizes. This regime permits a natural alternative: acquiring data using multiple AL strategies within a single batch. We refer to this as soft strategy selection and introduce FractAL, a method specifically designed for this task. FractAL infers a per-strategy reward using influence-function-based data attribution, which requires no additional retraining, and then computes budget shares for each strategy in the portfolio using online mirror descent. We benchmark FractAL across 7 setups spanning classification, regression, and genetic perturbation effect prediction. The results highlight that strategy selection is a hard problem: every existing method performs worse than random sampling on at least one setup. FractAL, however, matches or outperforms every baseline, including random sampling, on all 7 setups. Its allocations concentrate budget on the strongest strategies in the portfolio while pruning the weakest. FractAL is therefore a reliable choice for real-world deployments, where the optimal strategy is unknown in advance, an important step towards making AL practical for high-throughput experiments and modern scientific discovery.
发表机构
- New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。