arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IoUPD:用于多模态大语言模型视觉定位的IoU感知特权蒸馏

IoU-PD: IoU-Aware Privileged Distillation for Visual Grounding with Multimodal Large Language Models

Xiuyuan Zhu, Ke Lu, Hao Wu, Siwen Jiao, Zijin Du, Dongming Zhang, Jian Xue

arXiv 2607.15732首次发表:更新:

发表机构

University of Chinese Academy of Sciences; State Key Laboratory of Communication Content Cognition; Peng Cheng Laboratory(中国科学院大学; 通信内容认知技术国家重点实验室; 鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究多模态大语言模型视觉定位中训练与评估不匹配问题,提出IoUPD方法,利用真实框作特权指导,训练时学生模型接收原始信息,教师模型接收增强提示,经监督微调与特权蒸馏损失训练,推理时无需额外模块,实验显示该方法有改进。

AI 中文摘要

多模态大语言模型的视觉定位通常被表述为自回归坐标生成,即模型根据图像和指代表达提示输出作为文本的边界框坐标。虽然此接口简单且与指令跟随兼容,但它引入了训练和评估之间的不匹配:训练优化坐标字符串上的令牌级似然,而定位质量通过几何重叠来衡量。我们提出了IoUPD,一种用于坐标生成多模态大语言模型的IoU感知特权蒸馏方法。IoUPD不仅将真实框用作坐标目标,还用作训练时的特权指导。在训练期间,学生模型接收原始图像和提示,而冻结的教师模型接收带有框标记的图像和指示标记区域的增强提示。学生模型通过监督微调锚点和特权蒸馏损失进行训练,其令牌权重反映几何重要性和教师可靠性。在推理时,IoUPD不需要框覆盖、特权提示、教师分支或额外的预测模块。在标准指代表达定位基准上的实验表明,与强大的坐标生成基线相比,区域级有持续改进证据表明真实框除了作为坐标标签之外,还可以提供有用的特权指导。

英文摘要

Visual grounding with multimodal large language models is commonly formulated as autoregressive coordinate generation, where a model outputs bounding-box coordinates as text given an image and a referring-expression prompt. While this interface is simple and compatible with instruction following, it introduces a mismatch between training and evaluation: training optimizes token-level likelihood over coordinate strings, whereas grounding quality is measured by geometric overlap. We propose IoU-PD, an IoU-aware privileged distillation method for coordinate-generating multimodal large language models. IoU-PD uses ground-truth boxes not only as coordinate targets, but also as privileged training-time guidance. During training, the student receives the original image and prompt, while a frozen teacher receives a box-marked image and an augmented prompt that indicates the marked region. The student is trained with a supervised fine-tuning anchor and a privileged distillation loss whose token weights reflect both geometric importance and teacher reliability. At inference time, IoU-PD requires no box overlay, privileged hint, teacher branch, or additional prediction module. Experiments on standard referring-expression grounding benchmarks show consistent region-level improvements over strong coordinate-generating baselines, demonstrating that ground-truth boxes can provide useful privileged guidance beyond serving as coordinate labels. Project page: https://xyzzzh.github.io/IoU-PD/

Comments16 pages, 7 figures, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑