超越领域级自适应:面向个性化联邦视觉-语言模型的边际导向语义-外观交互校正
Beyond Domain-Level Adaptation: Margin-Oriented Semantic-Appearance Interaction Correction for Personalized Federated Vision-Language Models
AI总结:
针对联邦视觉-语言模型个性化中领域异质性问题,提出MOSAIC方法,通过决策感知危害性评分与低秩残差适配器建模细粒度类别-领域交互,并在多个基准上持续提升宏客户端top-1准确率。
AI中文摘要:
联邦参数高效微调使分布式客户端无需共享原始数据或更新完整骨干网络即可适配预训练的视觉-语言模型。然而,其有效性受到客户端间领域异质性的限制。现有的个性化方法将全局共享知识与客户端特定风格分离,但它们大多将每个领域视为与类别无关的变换。我们表明这种抽象是不充分的:与固定领域相关的跨领域位移在不同语义类别间存在差异,且仅这些类别-领域残差中的一部分会损害图像-文本决策边际。因此,我们提出边际导向语义-外观交互校正(MOSAIC),该方法首先构建一个决策感知的危害性评分,用于衡量训练得到的类别-领域残差是否使竞争性文本原型优于真实类别。随后,它通过低秩残差适配器建模细粒度的类别-领域交互,其中类别因子和残差基全局共享,而领域因子保持客户端私有。一个图像条件门进一步控制候选级校正,且有害对感知的重加权在局部优化期间优先处理决策相关残差。在Office31、OfficeHome和DomainNet100上的大量实验表明,MOSAIC在所有评估的领域偏移和联合领域-标签偏移设置中持续提升宏客户端top-1准确率。
英文摘要:
Federated parameter-efficient fine-tuning enables distributed clients to adapt pretrained vision-language models without sharing raw data or updating the full backbone. Its effectiveness, however, is limited by domain heterogeneity across clients. Existing personalized methods separate globally shared knowledge from client-specific style, but they largely treat each domain as a class-agnostic transformation. We show that this abstraction is insufficient: the cross-domain displacement associated with a fixed domain varies across semantic classes, and only a subset of these class-domain residuals damages the image-text decision margin. We therefore propose Margin-Oriented Semantic-Appearance Interaction Correction (MOSAIC), which first constructs a decision-aware harmfulness score that measures whether a training-derived class-domain residual favors a competing text prototype over the true class. It then models fine-grained class-domain interactions with a low-rank residual adapter whose class factors and residual basis are globally shared while domain factors remain client-private. An image-conditioned gate further controls candidate-wise correction, and harmful-pair-aware reweighting prioritizes decision-relevant residuals during local optimization. Extensive experiments on Office31, OfficeHome, and DomainNet100 demonstrate that MOSAIC consistently improves macro-client top-1 accuracy across all evaluated domain-shift and joint domain-label-shift settings.