arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10985cs.CV

MED-DSLC:通过域监督和对数校准进行多专家域分类

MED-DSLC: Multi-Expert-Domain Classification via Domain Supervision and Logit Calibration

  • University of California, San Diego(加利福尼亚大学圣地亚哥分校)

机构由 AI 辅助整理,请以论文原文为准。

Zheng Zeng, Deepak Sridhar, Nuno Vasconcelos

AI总结:

研究针对视觉语言模型在多域零样本分类中存在的问题,提出MED-DSLC方法,结合域监督训练和域内对数缩放,恢复全局对数可比性,是轻量级方案,经实验验证可提升准确率、鲁棒性和可扩展性。

AI中文摘要:

视觉语言模型(如CLIP)通过在共享嵌入空间中比较图像特征与文本提示来实现零样本分类,其基础是对数在任意候选类间的全局可比性。然而,使用LoRA等技术使VLMs适应细粒度域时,域内精度提高但域外精度下降,导致模型生态碎片化。多专家域分类试图通过合并在特定域上独立训练的LoRA来解决此问题,但独立训练导致域专家不再产生全局校准的对数,评估时会引发跨域干扰和预测错误。本文将域监督和跨域对数校准错误识别为可扩展多域零样本识别的关键问题,提出MED-DSLC,结合域监督训练和逐域对数缩放,以明确恢复全局对数可比性。MED-DSLC是一种轻量级的MED分类解决方案,能在减少跨域对数干扰的同时保持域内判别能力,实验表明其显著提高了平均准确率、跨域鲁棒性和可扩展性。结果表明,在高度数据不平衡设置下恢复输出级校准对于实现多域专业化下真正的零样本VLM至关重要。

英文摘要:

Vision-language models (VLMs) such as CLIP enable zero-shot classification by comparing image features with text prompts in a shared embedding space. A fundamental property underlying this capability is the global comparability of logits across arbitrary candidate classes. However, VLMs are often adapted to fine-grained domains using techniques such as LoRA. While this improves in-domain accuracy, out-of-domain accuracy degrades. This leads to a highly fragmented model ecosystem, with thousands of specialized models. Multi-Expert-Domain classification seeks to address this problem, by merging LoRAs trained independently on specialized domains. However, due to the independent training, the various domain experts no longer produce globally calibrated logits. As a result, when evaluating over the union of multiple domain-specific class sets, heterogeneous logit scales induce cross-domain interference and artificially high confidence for out-of-domain classes, inducing prediction errors. In this work, we identify domain supervision and cross-domain logit miscalibration as the key issue to scalable multi-domain zero-shot recognition. We propose MED-DSLC, combining domain supervised training and domain-wise logit scaling, to explicitly restore global logit comparability. MED-DSLC is a lightweight solution for MED classification, which is shown to preserve within-domain discrimination while reducing cross-domain logit interference with minimal data. Extensive experiments across diverse fine-grained benchmarks demonstrate that it substantially improves mean accuracy (+15\%), cross-domain robustness, and scalability in the size of MED classification problem. Our results show that restoring output-level calibration is essential under highly data imbalanced settings for achieving a truly zero-shot VLM under multi-domain specialization.

补充信息

↑