动量耦合的评分标准自适应用于细粒度图像描述生成
Momentum-Coupled Rubric Adaptation for Detailed Image Captioning
浏览论文内容
中文总结 AI 辅助
提出MoCo Rubric两阶段框架,通过共享参数多任务微调统一描述策略、评分标准生成与评判,并用动量模型平滑跟踪策略更新,在五个基准上实现72.83%平均胜率。
中文摘要 AI 辅助
细粒度图像描述生成需要对细粒度视觉内容进行准确且全面的描述,然而描述质量涉及事实准确性、信息覆盖度和清晰度等多个方面。与主要依赖高质量监督或整体奖励的传统方法相比,基于评分标准的强化学习将这些要求分解为明确的标准,并提供有针对性的、结构化的反馈。然而,现有方法通常使用独立的模型分别进行描述生成、评分标准构建和评判,这可能导致不同角色之间的解释不一致。一些动态评分标准方法在保持评判器固定的同时,交替更新描述策略和评分标准生成器,但分阶段优化仍可能使评分标准构建和评判与策略优化脱节。我们提出了MoCo Rubric,一个协调这些角色的两阶段框架。首先,基于角色条件、共享参数的多任务监督微调使单个视觉-语言模型能够同时充当描述策略、评分标准生成器和评分标准评判器。然后,生成器根据当前策略采样的描述、参考描述和图像证据在线构建评分标准。评判器提供基于评分标准的奖励,只有策略接收GRPO更新。由于策略更新会改变被评估的候选描述,我们使用策略参数的指数移动平均来更新一个由生成器和评判器共享的动量模型。这种渐进式转移使得两个评分标准角色无需单独的强化学习优化即可跟踪策略更新,同时平滑了在直接同步下可能破坏其评分标准能力的参数变化。在五个描述生成基准上,MoCo Rubric实现了72.83%的平均成对胜率,在盲排中取得最佳平均排名,并在基于描述的问题回答中取得最高平均得分。
英文摘要
Detailed image captioning requires accurate and comprehensive descriptions of fine-grained visual content, yet caption quality spans factual accuracy, information coverage, and clarity. Compared with conventional methods that rely mainly on high-quality supervision or holistic rewards, rubric-based reinforcement learning decomposes these requirements into explicit criteria and provides targeted, structured feedback. However, existing methods often use separate models for caption generation, rubric construction, and judging, which may lead to inconsistent interpretations across roles. Some dynamic rubric methods alternate updates between the caption policy and rubric generator while keeping the judge fixed, but staged optimization may still leave rubric construction and judging out of step with policy optimization. We propose MoCo Rubric, a two-stage framework that coordinates these roles. First, role-conditioned, shared-parameter multi-task supervised fine-tuning equips a single vision--language model to serve as the Caption Policy, Rubric Generator, and Rubric Judge. Then, the Generator constructs rubrics online from captions sampled by the current Policy, reference captions, and image evidence. The Judge provides rubric-based rewards, and only the Policy receives GRPO updates. As Policy updates change the candidates being evaluated, we use an exponential moving average of the Policy parameters to update one momentum model shared by the Generator and Judge. This gradual transfer lets both rubric roles track Policy updates without separate RL optimization while smoothing parameter changes that could disrupt their rubric capabilities under direct synchronization. Across five captioning benchmarks, MoCo Rubric achieves an average pairwise win rate of 72.83\%, the best mean rank in blind ranking, and the highest average score in caption-based question answering.
发表机构
- Xiaohongshu Inc.(小红书公司)
- the Joint SDU-NTU Centre for Artificial Intelligence Research (C-FAIR), Shandong University(山东大学-南洋理工大学联合人工智能研究中心(C-FAIR))
机构由 AI 辅助整理,请以论文原文为准。