arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ContourVLA:用于广义指代表达分割的闭环感知-动作轮廓策略

ContourVLA: A Closed-Loop Perception-Action Contour Policy for Generalized Referring Expression Segmentation

Ruicheng Zhang, Kaiwen Shen, Jiaqi Hou, Shuhan Yang, Junchao Huang, Kewei Zhang, Jun Zhou, Li Jiang, Shen Zhao

arXiv 2610.12107首次发表:更新:

发表机构

Sun Yat-sen University; Tsinghua University; The Chinese University of Hong Kong, Shenzhen; Peking University(中山大学; 清华大学; 香港中文大学(深圳); 北京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ContourVLA将广义指代表达分割转化为闭环视觉运动过程,通过EASS与DECT-GRPO优化,在多个基准数据集上显著提升了gIoU与mIoU性能。

AI 中文摘要

广义指代表达分割(GRES)需要动态平衡高层语义(用于识别语言指定的可变数量指代对象)与用于精确边界描绘的细粒度视觉证据。这一需求对现有级联视觉-语言架构构成挑战,这类架构通常依赖静态特征接口和单次掩码预测,限制了自适应感知与几何校正。我们提出ContourVLA,一种视觉-语言-动作策略,将GRES重新定义为闭环视觉运动过程,其中可编辑轮廓作为显式策略状态,调节多模态感知并通过几何动作块更新。进化感知语义调度(EASS)将轮廓引导的双向边界采样与多级别多模态特征的状态条件路由相结合,使感知适配每个轮廓状态。经监督初始化后,垃圾桶增强熵信用传输GRPO(DECT-GRPO)结合实例级信用优化离散 grounding 与连续轮廓动作,其部署奖励和信用源于考虑假阳性与漏检的软预测-目标对应关系。ContourVLA在gRefCOCO验证集、测试集A、测试集B上的gIoU较最强评估基线分别提升8.7、2.8、2.7个百分点,并在RefCOCO、RefCOCO+和RefCOCOg的全部8个划分中达到最高mIoU。

英文摘要

Generalized referring expression segmentation (GRES) requires dynamically balancing high-level semantics for identifying a variable number of language-specified referents with fine-grained visual evidence for precise boundary delineation. This requirement challenges existing cascaded vision-language architectures, which typically rely on static feature interfaces and single-pass mask prediction, limiting adaptive perception and geometric correction. We introduce ContourVLA, a vision-language-action policy that recasts GRES as a closed-loop visuomotor process, in which editable contours serve as explicit policy states that condition multimodal perception and are updated by geometric action chunks. Evolution-Aware Semantic Scheduling (EASS) couples contour-guided bidirectional boundary sampling with state-conditioned routing of multilevel multimodal features, adapting perception to each contour state. Following supervised initialization, Dustbin-Augmented Entropic Credit Transport GRPO (DECT-GRPO) jointly optimizes discrete grounding and continuous contour actions with instance-level credits. Its rollout rewards and credits are derived from soft prediction-target correspondences that account for false positives and missed targets. ContourVLA improves gIoU over the strongest evaluated baselines by 8.7, 2.8, and 2.7 points on gRefCOCO val, testA, and testB, respectively, and achieves the highest mIoU across all eight RefCOCO, RefCOCO+, and RefCOCOg splits.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑