发表机构
Carnegie Mellon University; Apple(卡内基梅隆大学; 苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对域外少样本检测,提出软提示方法,在跨模态边界放置少量可学习标记并初始化自空空间,以极小参数匹配LoRA精度且无遗忘,并扩展到操作任务。
AI 中文摘要
我们解决了视觉-语言模型(VLMs)在域外设置(如航空、工业和医学图像)中的少样本目标检测问题,仅使用十张标注图像进行监督。现有的适应方法是离散提示优化和LoRA微调。我们重新审视了第三种选择:软提示,其中优化少量连续提示标记,而预训练骨干保持冻结。我们确定了两个关键设计选择。首先,将提示标记放置在视觉和文本标记之间的跨模态边界处优于其他位置(10.0 vs. 8.4 mAP)。其次,从空空间标记初始化提示优于语义和随机初始化。通过这些选择,一到三个学习标记(平均7,168个参数)在Roboflow20-VL上匹配最佳LoRA配置(14.2 mAP,10-shot),同时训练参数减少超过20,000倍。软提示仍然更难优化,在随机种子之间表现出更高的方差。然而,与LoRA不同,它不会导致遗忘:与我们精度匹配的LoRA秩将NaturalBench VQA精度相对降低35%,在最大秩时升至56%,而软提示则保持预训练性能不变。学习到的标记行为类似于提示而非权重。它们可以迁移到更新的模型而无需重新训练(在Qwen3.5-9B上+0.8 mAP),并且可以转化为可读的提示,与提示搜索方法竞争(匹配DetPO并优于GEPA)。该方法还扩展到检测之外。在RoboCasa操作任务中,冻结的π_{0.5}视觉-语言-动作策略受益于软提示,在三个任务中的两个上匹配LoRA基线,当标记放置在梯度瓶颈处时。这些结果表明,现代VLM已经编码了专业领域所需的大部分知识;挑战在于学会如何提问。
英文摘要
We address few-shot object detection with vision-language models (VLMs) in out-of-domain settings such as aerial, industrial, and medical imagery, using only ten annotated images for supervision. Existing adaptation methods are discrete prompt optimization and LoRA fine-tuning. We revisit a third option: soft prompting, where a small number of continuous prompt tokens are optimized while the pretrained backbone remains frozen. We identify two key design choices. First, placing prompt tokens at the cross-modal boundary between visual and text tokens outperforms other placements (10.0 vs. 8.4 mAP). Second, initializing prompts from the empty space token outperforms semantic and random initialization. With these choices, one to three learned tokens (7,168 parameters on average) match the best LoRA configuration on Roboflow20-VL (14.2 mAP, 10-shot) while training over 20,000x fewer parameters. Soft prompting remains harder to optimize, exhibiting higher variance across random seeds. Unlike LoRA, however, it causes no forgetting: the LoRA rank matching our accuracy reduces NaturalBench VQA accuracy by 35% relative, rising to 56% at the largest rank, whereas soft prompting leaves pretrained performance unchanged. The learned tokens behave like prompts rather than weights. They transfer to a newer model without retraining (+0.8 mAP on Qwen3.5-9B) and can be verbalized into readable prompts competitive with prompt-search methods (matching DetPO and outperforming GEPA). The approach also extends beyond detection. On RoboCasa manipulation tasks, the frozen $π_{0.5}$ vision-language-action policy benefits from soft prompting, matching the LoRA baseline on two of three tasks when tokens are placed at the gradient bottleneck. These results suggest modern VLMs already encode much of what is needed for specialized domains; the challenge is learning how to ask.