arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

视觉-语言模型微调过程中的语义能力获取与特化

Semantic Capability Acquisition and Specialization During Vision-Language Model Fine-Tuning

Suguru Onda, Matthew Bailey, Ryan Farrell

arXiv 2610.07385首次发表:更新:

发表机构

Brigham Young University; Carnegie Mellon University(杨百翰大学; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于轨迹的框架和结构化语义路由,分析VLM微调中语义能力的获取、峰值与特化动态,发现不同能力峰值阶段不同,后期特化可能降低可迁移语义能力。

AI 中文摘要

对视觉-语言模型(VLM)进行微调通常是在单一的下游检查点进行评估,这掩盖了某个语义能力是从未被获取,还是在早期出现并在特化过程中衰退的问题。我们探究语义能力是如何被获取的、何时达到峰值、迁移效果如何,以及在部署时还保留了什么。我们将这些动态过程视为语义能力轨迹,追踪训练过程中基于身份和属性的能力变化。我们构建了一个基于轨迹的框架,将能力获取、能力特定最优状态和后期特化区分开来,并引入结构化语义路由(SSR)来研究监督的表征方式如何影响所获取的能力。在跨越DFN、MetaCLIP和OpenAI CLIP的六个预训练骨干网络上,我们表明微调能够获取超越预训练状态的语义能力,包括在保留评估中观察到的增益。非结构化的名称-属性监督产生强大的名称-属性检索,但相对较弱的无名称属性画像检索,而SSR则产生显著更强的无名称属性画像检索,并通过随机名称分支丢弃进一步增强。不同能力可能在不同阶段达到峰值,因此通过目标类名检索选择的检查点不一定与可迁移的语义最优状态重合。持续优化因此可以在保持强目标类名检索的同时,减少先前获取的可迁移语义能力。在一项代表性诊断研究中,这种后期特化与跨模态语义可访问性降低一致,而大量仅基于图像的类别结构仍然可用。

英文摘要

Fine-tuning vision-language models (VLMs) is typically evaluated at a single downstream checkpoint, obscuring whether a semantic capability was never acquired or emerged earlier and later declined during specialization. We ask how semantic capabilities are acquired, when they peak, how well they transfer, and what remains at deployment. We study these dynamics as a semantic capability trajectory, tracking identity- and attribute-based capabilities over training. We formulate a trajectory-based framework that separates capability acquisition, capability-specific optima, and later specialization, and introduce Structured Semantic Routing (SSR) to study how the representation of supervision shapes what is acquired. Across six pretrained backbones spanning DFN, MetaCLIP, and OpenAI CLIP, we show that fine-tuning can acquire semantic capability beyond the pretrained state, including gains observed on held-out evaluations. Unstructured name-and-attribute supervision produces strong name-and-attribute retrieval with comparatively weak name-free attribute-profile retrieval, whereas SSR yields substantially stronger name-free attribute-profile retrieval and is further strengthened by stochastic name-branch dropout. Different capabilities can peak at different stages, so a checkpoint selected by target class-name retrieval need not coincide with a transferable semantic optimum. Continued optimization can therefore preserve strong target class-name retrieval while reducing previously acquired transferable semantic capability. In a representative diagnostic study, this late specialization is consistent with reduced cross-modal semantic accessibility while substantial image-only class structure remains available.

Comments67 pages, including supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑