发表机构
Amap, Alibaba Group(高德,阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态大语言模型结构化视觉感知中响应级监督的粒度不匹配问题,提出MCR-GRPO框架,通过边界框级边际贡献分配优化,在多个基准上取得了优于现有GRPO基线的最先进性能。
AI 中文摘要
多模态大语言模型(MLLMs)日益被期望解决需要视觉识别、语言到物体的绑定、物体基数保留以及精确定位和分割输出的结构化感知任务。然而,现有的组相对强化学习方法仅提供响应级别的监督,这与结构化多物体预测存在粒度不匹配:单个优势被广播到响应中的所有 token,未区分各个边界框的贡献。为解决该不匹配,我们提出MCR-GRPO,一种直接从每个采样响应中推导边界框级信用的边际贡献分配框架。具体而言,边际贡献奖励(MCR)通过留一法比较估计每个预测边界框的贡献,测量当从响应中移除该边界框时匹配集值的变化。经响应内归一化后,提升集值的记录获得正信用,冗余或有害的记录则被抑制。为使边际归因稳定且具有信息性,我们进一步引入连续匹配集值评估器,其整合了置换不变匹配、计数感知归一化和分级定位。MCR-GRPO将归一化后的边界框级边际优势映射到生成每个边界框的 token 跨度,在保留GRPO响应级比较的同时,实现结构化多物体定位的边界框感知优化。在REC、DOD、分割和计数基准上的实验表明,其优于现有基于GRPO的基线,达到了最先进的性能。
英文摘要
Multimodal Large Language Models (MLLMs) are increasingly expected to solve structured perception tasks that require visual recognition, language-to-object binding, object cardinality preservation, and precisely localized grounding and segmentation outputs. However, existing group-relative reinforcement learning methods provide only response-level supervision, creating a granularity mismatch for structured multi-object prediction: a single advantage is broadcast to all tokens in a response, without distinguishing individual box contributions. To address this mismatch, we propose MCR-GRPO, a marginal contribution assignment framework that derives box-level credit directly from each sampled response. Specifically, Marginal Contribution Reward (MCR) estimates each predicted box's contribution through a leave-one-out comparison, measuring how the matched set value changes when the box is removed from the response. After within-response normalization, records that improve the set value receive positive credit, while redundant or harmful ones are suppressed. To make marginal attribution stable and informative, we further introduce a Continuous Matched Set Value Evaluator that integrates permutation-invariant matching, count-aware normalization, and graded localization. MCR-GRPO maps normalized box-level marginal advantages to the token spans that generated each box, preserving GRPO's response-level comparison while enabling box-aware optimization of structured multi-object grounding. Experiments across REC, DOD, segmentation, and counting benchmarks show state-of-the-art performance over prior GRPO-based baselines.