arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11581cs.CV

作为自身评论家的智能体:通过循环组相对策略优化统一区域理解与定位

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO

Xin Zhang, Haochen Wang, Yikang Zhou, Zhuochen Wang, Xiangtai Li, Robby T. Tan

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对多模态大语言模型的区域理解与定位问题,提出循环组相对策略优化框架CycleGRPO,利用任务对偶性构建自我评估范式,仅需区域输入,通过质量感知奖励评估字幕,在多基准测试中提升能力,为推进MLLMs像素级能力提供新途径。

中文摘要 AI 辅助

本文介绍了作为自身评论家的智能体,一种统一的强化学习框架——循环组相对策略优化(CycleGRPO),用于联合优化多模态大语言模型(MLLMs)的区域理解与定位。与现有单独的管道不同,利用两项任务的内在对偶性构建了“区域→文本→区域”的自我评估强化学习范式。单个MLLM先作为智能体生成区域字幕,再转为评论家将生成的文本回归空间域。CycleGRPO仅需区域输入,通过质量感知令牌级循环一致性奖励评估文本字幕语义可辨别性。基于SAMTok构建的该框架在多个基准测试中同时提升了两种能力,无需特定任务微调,为提升MLLMs像素级能力提供了直接且可扩展的方法。代码和模型已发布。

英文摘要

This paper introduces Actor as Its Own Critic, a unified reinforcement learning framework, Cycle Group Relative Policy Optimization (CycleGRPO), that jointly optimizes region understanding and localization for Multimodal Large Language Models (MLLMs). Unlike existing separate pipelines, we leverage the inherent duality between the two tasks to construct a self-evaluating reinforcement learning paradigm: "region $\to$ text $\to$ region''. Specifically, a single MLLM first acts as the actor to generate region captions, then immediately transitions to a critic to ground its generated text back in the spatial domain. Therefore, CycleGRPO requires only region inputs, e.g., masks or bounding boxes, entirely bypassing the need for textual ground truths. A quality-aware token-level cycle-consistency reward is employed to assess the semantic discriminability of text captions via their physical localization accuracy. Empirically, built upon SAMTok, our CycleGRPO framework successfully bootstraps both capabilities simultaneously. Without any task-specific fine-tuning, the framework yields consistent performance gains across a wide range of benchmarks, including region captioning, region VQA, grounded dialogue, and referring segmentation. Overall, CycleGRPO offers a straightforward and scalable way to advance pixel-level capabilities in MLLMs. Code and models are released at https://github.com/devinxzhang/CycleGRPO.

发表机构

  • National University of Singapore(新加坡国立大学)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Nanyang Technological University(南洋理工大学)
  • Wuhan University(武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑