arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22323cs.CVcs.LG

ALPINE:面向参数与样本高效少样本学习的自适应定位

ALPINE: Adaptive Localization for Parameter- and Sample-Efficient Few-Shot Learning

Neeraj Yadav

首次发表
浏览论文内容

中文总结 AI 辅助

提出超轻量级空间关系架构ALPINE,以更少参数和样本在少样本图像分类中超越现有基线,并揭示内容自适应补丁定位器是性能关键。

中文摘要 AI 辅助

少样本学习研究主要仅以准确率进行评估,而对达到该准确率所需的参数和训练样本预算关注有限——这对于没有大规模计算资源的从业者而言是一个实际约束。我们提出了一种用于少样本图像分类的超轻量级(22,249-34,917参数)空间关系架构,该架构将固定的Gabor边缘能量引导与窗口化、内容自适应的补丁定位器相结合。在严格匹配的、等情节预算协议下(250个元训练情节、5个规范种子、每个种子600个评估情节),我们的架构在CIFAR-FS和MiniImageNet上,相对于原型网络、关系网络和MAML,在所有五个种子上均一致地实现了5-shot准确率提升,同时使用的参数比任何基线少27-53%。它还在更少的训练情节中收敛,更好地泛化到未见过的细粒度领域(CUB-200-2011鸟类,零重新训练),并且比所有三个基线对50%遮挡和25%空间平移更具鲁棒性。一系列证伪消融实验——在推理时将关系令牌置零以及完全不带它们重新训练——表明该架构的成对关系计算虽然存在,但并非其性能的主要驱动因素;内容自适应的补丁定位器才是。我们诚实地报告了这一结果,并附上容量扫描,显示在22-35k参数附近存在真实的准确率平台,同时发布完整的种子级结果和检查点哈希以供可复现性。

英文摘要

Few-shot learning research is predominantly evaluated on accuracy alone, with limited attention to the parameter and training-sample budgets required to reach that accuracy - a real constraint for practitioners without large-scale compute. We present an ultra-lightweight (22,249-34,917 parameter) spatial-relational architecture for few-shot image classification that combines fixed Gabor edge-energy guidance with a windowed, content-adaptive patch locator. Under a strictly matched, iso-episode-budget protocol (250 meta-training episodes, 5 canonical seeds, 600 evaluation episodes per seed), our architecture achieves 5-shot accuracy gains, consistent across all five seeds, over Prototypical Networks, Relation Networks, and MAML on both CIFAR-FS and MiniImageNet, while using 27-53% fewer parameters than any baseline. It also converges in fewer training episodes, generalizes better to an unseen fine-grained domain (CUB-200-2011 birds, zero retraining), and is more robust to 50% occlusion and 25% spatial translation than all three baselines. A series of falsification ablations - zeroing relational tokens at inference and retraining without them entirely - shows that the architecture's pairwise relational computation, while present, is not the primary driver of its performance; the content-adaptive patch locator is. We report this honestly, together with a capacity sweep showing a genuine accuracy plateau near 22-35k parameters, and release full seed-level results and checkpoint hashes for reproducibility.

补充信息

↑