发表机构
School of Computer Science and Engineering, Sun Yat-sen University(中山大学计算机科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对无源域自适应的源衍生证据遗忘问题,提出COSMO方法,通过锚定共享共识的协同适配实现样本级可靠性分配,在四个基准上达到最优性能,平衡了源证据保留与VLM互补证据吸收。
AI 中文摘要
无源域自适应(SFDA)是在隐私或存储约束下,将源训练模型适配至无标签目标域的实用设置,无需源数据。但其自生成的监督信号会在域偏移较大时强化源偏差。预训练视觉语言模型(VLM)提供互补语义知识,但源模型与VLM的相对可靠性在不同目标样本间存在差异。现有跨模型指导未明确考虑该差异,冲突时可能覆盖有效的源衍生证据,该缺陷被称为“源衍生证据遗忘”。我们将VLM引导的SFDA建模为样本级可靠性分配问题,提出共识驱动偏移调制(COSMO)。COSMO通过锚定共享共识的协同适配替代专家间指导,先形成偏好更集中预测的样本特定初始共识;适配过程中,COSMO重新聚合两个分支的演化证据,并基于共识不确定性和训练进度调节所得共识偏离初始锚点的程度,使共享监督既锚定又自适应。在四个基准测试中,COSMO在匹配VLM骨干下达到最优性能,进一步分析表明其能更好平衡有效源衍生证据的保留与互补VLM证据的吸收。
英文摘要
Source-free domain adaptation (SFDA) adapts a source-trained model to an unlabeled target domain without source data, a practical setting under privacy or storage constraints. Yet its self-generated supervision can reinforce source bias under substantial domain shifts. Pretrained vision-language models (VLMs) offer complementary semantic knowledge, but the relative reliability of the source model and VLM varies across target samples. Existing cross-model guidance does not explicitly account for this variation and may overwrite valid source-derived evidence under conflict, a failure we term source-derived evidence forgetting. We formulate VLM-guided SFDA as a sample-wise reliability-allocation problem and propose Consensus-Driven Shift Modulation (COSMO). COSMO replaces expert-to-expert guidance with co-adaptation through an anchored shared consensus. It first forms a sample-specific initial consensus that favors the more concentrated prediction. During adaptation, COSMO re-aggregates both branches' evolving evidence and regulates how far the resulting consensus moves from its initial anchor based on consensus uncertainty and training progress. This keeps the shared supervision anchored yet adaptive. Across four benchmarks, COSMO achieves state-of-the-art performance under matched VLM backbones. Further analyses indicate that it better balances the retention of valid source-derived evidence with the absorption of complementary VLM evidence.
Comments30 pages, 7 figures