AI 中文总结
该研究提出受生物启发的激活函数HAND,可减少图像分类任务的训练时间与数据量,提升样本效率,在ConvNeXt-tiny等模型及ImageNet等数据集上验证了其有效性。
AI 中文摘要
深度神经网络(DNNs)存在人类所未出现的鲁棒性与泛化问题,且数据效率远低于人类,需要大量训练样本才能准确分类新样本。归纳偏置可通过内置机制改善泛化,减少对数据学习的依赖。我们将受生物启发的归纳偏置融入新激活函数HAND(Homeostasis、加速非线性与除法归一化),并在图像分类任务的卷积神经网络(CNNs)上验证其有效性。使用HAND时,ConvNeXt-tiny模型仅需25个训练轮次(epoch)即可达到未修改模型在ImageNet1k数据集上200个轮次的准确率。与归纳偏置的效果一致,性能差距随训练时间增加而缩小,随数据增强强度提升而增大。当训练数据量减少且类别分布不均(长尾ImageNet)时,准确率提升幅度更大,且不会随训练时间增加而下降。HAND对常见损坏数据的泛化性能及未知类别的弃权(不执行)能力无负面影响或有所提升,结果可推广至多种CNN架构与训练数据集。因此,HAND可减少所需的训练时间和/或训练数据的数量与多样性,助力提升样本效率。
英文摘要
DNNs exhibit robustness and generalisation issues not seen in humans. They are also far less data-efficient learners, requiring considerably more training samples to accurately classify novel exemplars. Inductive bias could help with these issues by providing in-built mechanisms to improve generalisation, and hence, reduce reliance on learning from data. We incorporate a biologically-inspired inductive bias into a new activation function, HAND (Homeostasis, Accelerating Nonlinearity, and Divisive-nomalisation), and show its effectiveness with CNNs trained on image classification. Using HAND a ConvNeXt-tiny required 25 training epochs to reach the same accuracy on ImageNet1k as the unmodified model achieved after 200 epochs. Consistent with the effects of an inductive bias, the performance gap reduced with training time and increased data augmentation. When the volume of training data was reduced and unevenly distributed between classes (Long-tailed ImageNet) the improvements in accuracy were even larger and did not reduce with increased training time. Generalisation performance with the common-corruptions data, and the ability to reject samples from unknown classes, were unaffected or improved by HAND. Results generalised across CNN architectures and training data-sets. HAND can, therefore, reduce the required training time and/or the required volume and variety of training data, helping to improve sample efficiency.