arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VICAL:用于长尾视觉识别的邻域一致性对齐

VICAL: Vicinal Consistency Alignment for Long-Tailed Visual Recognition

Jiangang Zhu, Zheng Wang, Bin Zhu, Yi-Ping Phoebe Chen, Jingjing Chen

arXiv 2609.04948首次发表:更新:

发表机构

Fudan University; Zhejiang University of Technology; Singapore Management University; La Trobe University(复旦大学; 浙江工业大学; 新加坡管理大学; 拉筹伯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对长尾视觉识别,提出 VICAL 框架,通过自一致性学习和深度集成蒸馏减少预测方差,在多个长尾数据集上优于现有方法。

AI 中文摘要

多专家模型已成为长尾学习的主流范式,很大程度上归因于人们认为它们能从专家多样性中获益。然而,我们重新审视这一核心假设,发现由 logit 调整或显式正则化器诱导的多样性并不能保证更好的集成准确率。我们的研究表明,多专家模型从方差减少中获益比从多样性最大化中更多。我们引入 VICAL,这是一种邻域一致性对齐(VIcinal Consistency ALignment)框架,它不通过强制专家多样性,而是通过减少预测方差来改进长尾识别。具体而言,我们的方法包含两个关键组件:自一致性学习和深度集成蒸馏。自一致性学习阻止模型依赖不稳定的高频信息,平滑局部损失景观并缓解过拟合,尤其是针对尾部类别。深度集成蒸馏利用低分辨率视图促进跨专家的低频语义一致性,从而避免与已有知识的优化冲突。在 CIFAR-LT、ImageNet-LT 和 iNaturalist 2018 上的大量实验表明,VICAL 始终优于最先进的方法,验证了我们以一致性为驱动的设计的有效性。我们的代码可在 https://github.com/VICAL 获取。

英文摘要

Multi-expert models have become the dominant paradigm for long-tailed learning, largely attributed to their presumed ability to benefit from expert diversity. However, we revisit this central assumption and reveal that diversity induced by logit adjustment or explicit regularizers does not guarantee better ensemble accuracy. Our work suggests that multi-expert models benefit more from variance reduction than diversity maximization. We introduce \textbf{VICAL}, a \textbf{VI}cinal \textbf{C}onsistency \textbf{AL}ignment framework that improves long-tailed recognition not by enforcing expert diversity, but by reducing prediction variance. Specifically, our approach comprises two key components: Self-Consistency Learning and Deep Ensemble Distillation. Self-Consistency Learning discourages reliance on unstable high-frequency information, smoothing the local loss landscape and mitigating overfitting, especially for tail classes. Deep Ensemble Distillation promotes cross-expert low-frequency semantic agreement using a low-resolution view, thereby sidestepping optimization conflicts with established knowledge. Extensive experiments on CIFAR-LT, ImageNet-LT, and iNaturalist 2018 show that VICAL consistently outperforms state-of-the-art methods, validating the effectiveness of our consistency-driven design. Our code is available at \href{https://github.com/FlamieZhu/Vicinal-Consistency-Alignment}{VICAL}.

CommentsAccepted to ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑