arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CoEvolve:基于双向状态细化的构造到编辑视觉定位

CoEvolve: Construct-to-Edit Visual Grounding with Bidirectional State Refinement

Dongwei Sun, Yujie Zhang, Bowen Yao, Pei Liu, Jing Yao, Xiangyong Cao

arXiv 2610.01710首次发表:更新:

发表机构

Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

CoEvolve提出构造到编辑框架,通过显式状态构造和双向细化分离视觉定位,在9B参数下媲美241B模型,单次细化提升边界框重叠率超27个百分点。

AI 中文摘要

视觉定位通过边界框定位语言所描述的对象。大多数多模态定位模型将目标识别、空间推理和边界估计压缩为一次终端预测。自由形式的推理使推理过程在语言上明确,但不一定暴露可测量、可编辑的空间状态。因此,中间定位错误难以诊断和纠正,导致错误的区域选择和模糊的边界在最终边界框中持续存在。我们引入了CoEvolve,一个构造到编辑框架,将定位分离为显式状态构造和状态编辑。区域演化强化(RER)将定位分析组织为渐进式的语义-空间轨迹,每个推理步骤都承诺一个显式候选区域。双向去噪细化器(BDR)将推理文本视为固定语义上下文,通过双向同位重建细化轨迹的坐标字段。几何和行为层面的目标提供目标几何和编辑偏好信号,用于巩固可靠候选、保留准确输入或向注释方向修正。评估涵盖自然图像和遥感定位。使用9B骨干,CoEvolve在定位精度上可与高达241B参数的模型媲美。在受控损坏下,单次BDR传递将平均边界框重叠率提高超过27个百分点,展示了从重大定位错误中强恢复的能力。状态源比较进一步支持显式状态构造与源匹配编辑的互补性。项目位于此https URL。

英文摘要

Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reasoning, and boundary estimation into one terminal prediction. Free-form rationales make reasoning linguistically explicit but do not necessarily expose measurable, editable spatial states. Intermediate localization errors are therefore difficult to diagnose and correct, allowing incorrect region choices and imprecise boundaries to persist in the final box. We introduce CoEvolve, a construct-to-edit framework that separates grounding into explicit state construction and state editing. Region-Evolution Reinforcement (RER) organizes grounding analysis into a progressive semantic--spatial trajectory, with each reasoning step committing to an explicit candidate region. Bidirectional Denoising Refiner (BDR) treats the reasoning text as fixed semantic context and refines the trajectory's coordinate fields through bidirectional same-position reconstruction. Geometry- and behavior-level objectives provide target geometry and edit-preference signals for consolidating reliable candidates, preserving accurate inputs, or correcting toward annotations. Evaluations cover natural-image and remote-sensing grounding. With a 9B backbone, CoEvolve rivals models up to 241B parameters in grounding accuracy. Under controlled corruption, a single BDR pass improves mean box overlap by over 27 percentage points, demonstrating strong recovery from substantial localization errors. State-source comparisons further support the complementarity of explicit state construction and source-matched editing. The project is at https://sundongwei.github.io/CoEvolve_Project/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑