发表机构
AI Laboratory, University Metropolitan Tirana; Laboratory of Images Signals and Intelligent Systems, ESIEE Paris(地拉那都市大学人工智能实验室; 巴黎高等电子与电工工程师学校图像信号与智能系统实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对CNN训练中小批量随机打乱导致的低效问题,提出受A*启发的批量选择方法,通过特定分数排序批次,在MedMNIST-v2基准任务中,该方法在部分任务上超越ResNet基线,训练更快,证明智能批量排序可弥补架构不足。
AI 中文摘要
训练卷积神经网络(CNN)时常用随机打乱的小批量数据,这存在收敛慢和学习信号减弱的问题。我们提出受A*启发的批量选择(A*-BS)来解决这些低效问题,它将小批量调度视为启发式搜索问题,通过结合基于损失的难度度量和重用惩罚的A*类分数对批次进行排序。在MedMNIST-v2基准的十二个二维分类任务上评估A*-BS,使用约2.25×10^5参数的简单架构,与基准报告的ResNet-18和ResNet-50基线比较。结果表明,A*-BS在一半任务上比两个ResNet基线有更高准确率和AUC,消融实验显示其在所有任务上优于随机批量打乱,且训练速度更快。这表明智能批量排序可部分弥补架构复杂性降低的问题,为更深模型提供计算高效的替代方案。
英文摘要
Common practice when training Convolutional Neural Networks (CNNs) is to use randomly shuffled mini-batches. This creates two limitations: slower convergence, and a diminishing learning signal, since many samples are quickly classified as easy during training. We address these inefficiencies with A*-Inspired Batch Selection (A*-BS), a lightweight, model-agnostic strategy that formulates mini-batch scheduling as a heuristic search problem. Each batch is treated as a node in a search space and ranked using an A*-like score combining a loss-based difficulty measure with a reuse penalty. This encourages informative gradient updates and batch diversity throughout training, without modifying network architectures or optimization algorithms, so it integrates seamlessly into existing pipelines. We evaluate A*-BS on the twelve 2D classification tasks of the MedMNIST-v2 benchmark, using a deliberately simple architecture of approximately 2.25x10^5 parameters, compared against the ResNet-18 and ResNet-50 baselines reported by the benchmark. On half of these tasks, the lightweight model with A*-BS reaches higher accuracy and AUC than both ResNet baselines, with relative gains of up to 15%. An ablation under identical architecture and hyperparameters shows A*-BS outperforms random batch shuffling on all twelve tasks. Wall-clock measurements further show the lightweight CNN with A*-BS trains substantially faster than ResNet-18 and ResNet-50 on identical hardware. These results indicate that intelligent batch ordering can partially compensate for reduced architectural complexity, offering a computationally efficient alternative to deeper models, with reliability reinforced by strong performance even against deeper, more sophisticated architectures.