arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PETR:面向视觉-语言模型的免训练路由提示集成

PETR: Prompt Ensembling with Training-free Routing for Vision-Language Models

Weihan Cai, Hao Tan, Xinping Gao, Shibiao Xu, Jun Wan

arXiv 2609.23600首次发表:更新:

发表机构

State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences; School of Artificial Intelligence, University of Chinese Academy of Sciences; Purple Mountain Laboratories; Beijing University of Posts and Telecommunications(中国科学院自动化研究所多模态人工智能系统国家重点实验室; 中国科学院大学人工智能学院; 紫金山实验室; 北京邮电大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PETR提出双提示架构与免训练统计路由,在保持已见类别性能的同时提升未见类别泛化,在11个基准上取得新最先进结果。

AI 中文摘要

提示学习能高效地将视觉-语言模型(VLMs)适应到下游任务,但在已见类别上的增益往往以牺牲对未见类别的泛化为代价。为解决这一局限,我们提出了免训练路由的提示集成(PETR),其关键创新在于精心设计的双提示架构:两个互补的提示分别从不同的数据和目标中学习,以分别强调已见类别的判别和未见类别的泛化。在训练期间,两个提示均使用共享的冻结CLIP骨干进行微调,并从训练集logits中收集统计信息。在推理时,我们确定每个测试样本与已见数据的相似度,并将样本路由到最合适的提示分支。据我们所知,这是首个基于统计相似性执行免训练自适应路由的提示调优框架。该设计提供了可解释的路由信号,并避免了常见的MoE式路由病态,如路由器训练不稳定和负载不平衡。在11个基准数据集上的大量实验表明,我们的框架在已见和未见类别上均持续优于先前方法,取得了新的最先进结果。

英文摘要

Prompt learning efficiently adapts vision-language models (VLMs) to downstream tasks, but gains on seen classes often come at the expense of generalization to unseen classes. To address this limitation, we propose prompt ensembling with training-free routing (PETR), whose key innovation is a carefully designed dual-prompt architecture: two complementary prompts are learned from different data and objectives to emphasize seen class discrimination and unseen-class generalization, respectively. During training, both prompts are fine-tuned using a shared frozen CLIP backbone, and statistical information is collected from the training set logits. At inference time, we determine the similarity of each test sample to seen data, and route the sample to the most appropriate prompt branch. To the best of our knowledge, this is the first prompt tuning framework that performs training-free adaptive routing based on statistical similarity. This design provides an interpretable routing signal and avoids common MoE-style routing pathologies, such as router training instability and load imbalance. Extensive experiments on 11 benchmark datasets demonstrate that our framework consistently outperforms previous methods on both seen and unseen classes, achieving new state-of-the-art results.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑