发表机构
KT Corporation; Pohang University of Science and Technology (POSTECH); National AI Research Lab(KT公司; 浦项科技大学; 国家人工智能研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出文本提示提升(TPB)框架,通过AdaBoost策略将每个文本提示分类器视为弱学习器,顺序集成以聚焦难分类样本,提升少样本分类精度并实现跨模型迁移。
AI 中文摘要
预训练的视觉-语言模型(VLM)的分类准确性依赖于文本提示的质量。手工设计的模板和大语言模型(LLM)生成的描述不仅使预测更具可解释性,还允许在不同VLM之间重用相同的提示。最近的工作利用少量标注图像构建任务适应的文本提示。然而,现有的少样本文本提示方法在提示构建过程中并未明确关注误分类样本,导致即使增加更多样本,改进也仅微乎其微。为了充分利用少样本监督,我们提出文本提示提升(TPB),一种受AdaBoost启发的框架,将每个基于文本提示的分类器视为弱学习器,并通过明确针对困难的、误分类的样本,顺序地将它们聚合为一个强集成。大量实验表明,TPB在文本空间中保留了任务固有的、模型无关的线索,实现了鲁棒的跨模型迁移。在十一个分类基准上,TPB提高了源模型的准确性,并在迁移到更大、更强大的VLM时保持了样本驱动的改进,而现有方法难以维持这种改进。
英文摘要
The classification accuracy of pretrained Vision-Language Models (VLMs) relies on the quality of the text prompts. Handcrafted templates and Large Language Model (LLM)-generated descriptions not only make predictions more interpretable, but also enable reuse of the same prompts across heterogeneous VLMs. Recent works construct task-adapted text prompts with a small number of labeled images. However, existing few-shot text prompting methods do not explicitly focus on misclassified examples during prompt construction, leading to only marginal improvements even as more shots become available. To fully exploit few-shot supervision, we propose Text Prompt Boosting (TPB), an AdaBoost-inspired framework that treats each text-prompt-based classifier as a weak learner and sequentially aggregates them into a strong ensemble by explicitly targeting hard, misclassified examples. Extensive experiments show that TPB preserves task-intrinsic, model-agnostic cues in text space, enabling robust cross-model transfer. Across eleven classification benchmarks, TPB improves accuracy on the source model and preserves shot-driven gains when transferred to larger, more capable VLMs, where existing methods struggle to sustain such improvements.
CommentsAccepted to ECCV 2026 Spotlight. Minor typo correction in the ECCV camera-ready