arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉语言模型的无监督长尾适应

Unsupervised Long-Tailed Adaptation of Vision-Language Models

Keliang Chen, Yaxin Hou, Hui Liu, Yuheng Jia

arXiv 2610.07903首次发表:更新:

发表机构

Southeast University; Saint Francis University(东南大学; 圣方济各大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对无标签数据长尾分布导致视觉语言模型头部类性能下降的问题,提出MARS模型,通过边界保留对齐和边际感知自精炼,在九个基准上平均准确率提升4.71个百分点。

AI 中文摘要

通过利用从无标签数据生成的伪标签,视觉语言模型适应下游任务已取得显著成功。现有方法通常假设无标签数据分布是均匀的,因此生成的伪标签分布也是均匀的。然而,现实世界的数据分布往往是长尾的。为解决这一问题,我们形式化了一个新场景,称为无监督长尾适应(ULTA)。在此场景下,现有方法表现出一种对比现象:头部类性能急剧下降,这与监督长尾学习中尾部类受损最严重的情况截然不同。特别是,我们发现分布不匹配不仅侵蚀了头部类的边界,还将头部样本推入易混淆的类别,加剧了模型固有的偏差。为解决这些问题,我们提出了一种名为“边界保留对齐与结构对齐的边际感知精炼”(MARS)的新模型。具体而言,我们通过边界保留对齐来缓解头部类边界侵蚀,该方法将零样本视觉语言模型作为固定的视觉参考,以抑制训练目标中缺乏视觉支持的概率增加。在此基础上,我们引入了边际感知自精炼,采用动态调整策略来精炼尾部类和易混淆类,同时防止预测偏差。在九个基准数据集上的大量实验表明,MARS优于最先进的方法,平均准确率提升了4.71个百分点。

英文摘要

Adapting vision-language models to downstream tasks has achieved remarkable success by leveraging pseudo-labels generated from unlabeled data. Existing methods typically assume a uniform unlabeled data distribution, and thus the resulting pseudo-label distribution is likewise uniform. However, real-world data distributions are often long-tailed. To tackle this, we formalize a new scenario termed Unsupervised Long-Tailed Adaptation (ULTA). Under this scenario, existing methods exhibit a contrasting phenomenon: head-class performance drops sharply, which is distinct from supervised long-tailed learning where tail classes suffer the most. In particular, we uncover that the distributional mismatch not only erodes head-class boundaries, but also pushes head samples into confusable classes, reinforcing the model's inherent bias. To address these issues, we propose a novel model called Margin-Aware Refinement with Structural alignment (MARS). Specifically, we mitigate head-class boundary erosion via Boundary-Preserving Alignment, which takes the zero-shot VLM as a fixed visual reference to suppress probability increases that lack visual support in the training targets. Building upon this, we introduce Margin-aware Self-Refinement, which employs a dynamic adjustment strategy to refine tail and confusable classes while preventing prediction bias. Extensive experiments on nine benchmark datasets demonstrate that MARS outperforms state-of-the-art methods, achieving an average accuracy improvement of 4.71 percentage points.

Comments18 pages, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑