指令蒸馏:将文本指令作为视觉示例
Instruction Distillation: Text Instructions as Visual Examples
浏览论文内容
中文总结 AI 辅助
该研究针对视觉上下文学习的高成本问题,提出指令蒸馏方法,为每个训练图像生成结构化指令,在细粒度视觉分类任务中提升性能并降低推理成本,且视觉与文本ICL信号互补。
中文摘要 AI 辅助
采用多模态大语言模型(MLLM)的视觉上下文学习(ICL)对细粒度视觉分类有效,但每个检索到的图像示例会消耗数百个上下文token,导致大K设置在推理规模下成本过高。我们提出指令蒸馏:一种离线流程,其中MLLM为每个训练图像生成结构化识别指令,编码一般外观线索、与视觉相似类区分的特征以及常见混淆点。与以往为每个类生成单一描述的工作不同,我们的指令针对每个训练图像生成,保留了类内视觉多样性,而按类描述会丢失该多样性。推理时,我们研究了共享单个CLIP检索索引的五种配置:零样本、图像ICL、仅指令ICL,以及两种混合变体,其中检索到的邻居在图像和指令之间拆分。在七个细粒度基准和两个MLLM主干上,基于指令的流水线在K=1时与图像ICL相当或优于其性能,在K=5时每个查询的token减少2.9倍,推理延迟降低3.3倍。混合配置进一步表明,视觉和文本ICL信号具有互补性:图像提供学习和观察的视觉模式,指令提供明确的规则和逻辑。当同时提供这两者时,上下文质量会提升,这种提升在性能上表现明显。
英文摘要
Visual in-context learning (ICL) with multimodal large language models (MLLMs) is effective for fine-grained visual classification, but each retrieved image example consumes several hundred context tokens, making large-$K$ settings prohibitively expensive at inference scale. We propose Instruction Distillation: an offline procedure in which the MLLM itself generates, for each individual training image, a structured identification instruction encoding general appearance cues, features that differentiate the class from visually similar ones, and a common confusion point. Unlike prior work that produces a single description per class, our instructions are generated per training image, preserving the intra-class visual diversity that per-class descriptions collapse. At inference time, we study five configurations sharing a single CLIP retrieval index: zero-shot, image ICL, instruction-only ICL, and two hybrid variants in which retrieved neighbors are split between images and instructions. Across seven fine-grained benchmarks and two MLLM backbones, instruction based pipelines match, or exceeds image ICL at $K{=}1$ and reduces per-query tokens by $2.9\times$ and inference latency by $3.3\times$ at $K{=}5$. Hybrid configurations further show that visual and textual ICL signals are complementary, images give visual patterns to learn and see, while instructions give explicit rules and logic. When both of these are provided, the quality of context improves, which is noticeable in the performance.
发表机构
- IIT Kanpur(印度理工学院坎普尔分校)
- Adobe Research(奥多比研究院)
机构由 AI 辅助整理,请以论文原文为准。