AI 中文总结
针对MLLMs因训练缺负样本导致无法拒绝不存在对象的问题,提出RC-GRPO方法,在保持定位精度的同时增强拒绝能力,在三个GREC基准上表现优异。
AI 中文摘要
我们解决了具有挑战性但尚未充分探索的广义指代表达理解(GREC)任务,该任务要求模型在文本表达式描述的对象存在(正样本)时对其进行定位,在对象不存在(负样本)时输出弃权(不执行)。尽管多模态大语言模型(MLLMs)擅长定位存在的对象,但由于训练期间缺少负样本,它们往往无法对不存在的对象进行拒绝,从而产生幻觉的边界框。现有后训练方法如监督微调(SFT)和强化学习(RL)可增强拒绝行为,但通常会降低正样本的定位精度,损害模型的核心能力。为解决该问题,我们提出拒绝校准组相对策略优化(RC-GRPO),这是一种校准后的RL策略,可在保持定位性能的同时增强MLLMs的拒绝能力。该策略在 rollout 中强制负样本输出“None”以进行有效的优势估计,并施加惩罚以防止正样本上的过度拒绝,在精度和可靠性之间实现平衡权衡。第二阶段推理强化进一步巩固了因果理解和可解释性。在三个GREC基准上的实验表明,RC-GRPO在保持强拒绝能力的同时实现了优异的定位精度。
英文摘要
We tackle the challenging yet underexplored task of Generalized Referring Expression Comprehension (GREC), which requires a model to localize the object described by a textual expression when it exists (positive sample) and to refuse output when it does not (negative sample). Although Multimodal Large Language Models (MLLMs) excel at localizing existing objects, they often fail to reject nonexistent ones due to the absence of negative samples during training, producing hallucinated bounding boxes. Existing post-training approaches such as supervised fine-tuning (SFT) and reinforcement learning (RL) enhance refusal behavior but usually degrade localization accuracy on positive samples, undermining the model's core competence. To address this, we propose Refusal-Calibrated Group Relative Policy Optimization (RC-GRPO), a calibrated RL strategy that strengthens the refusal ability of MLLMs while preserving localization performance. It enforces "None" outputs in rollouts for valid advantage estimation on negative samples and applies a penalty to prevent over-refusal on positives, achieving a balanced trade-off between accuracy and reliability. A second-stage reasoning reinforcement further consolidates causal understanding and interpretability. Experiments on three GREC benchmarks demonstrate that RC-GRPO attains superior localization accuracy while maintaining strong refusal capability.