arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14148cs.CV

SCVIB:面向多轮个性化定位的可编辑状态条件视觉实例绑定

SCVIB: Editable State-Conditioned Visual Instance Binding forMulti-Turn Personalized Localization

Xiongtai Yang, Ziyan He, Tao Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出多轮个性化定位设置SCVIB,结合TT-VG与VEGA方法构建测试平台,解决多轮定位中支持证据利用不足问题,相关指标优于对比方法。

中文摘要 AI 辅助

我们提出了可编辑状态条件视觉实例绑定,这是一种多轮定位设置,其中在多轮中引入若干由支持项定义的实例,且由协议定义的状态事件决定最终目标。我们将该设置实例化为SCVIB,包含1050个经人工验证的支持-查询基础对,以及跨越5个视觉域、3个难度级别和4个目标状态依赖组的1500个 episodes(场景)。直接无序列推理仅达到60.13%的Joint@0.5指标,表明解析最终参考并不能确保有效利用相应视觉证据进行查询侧定位。我们通过TT-VG(Transition-Tree Visual Grounding,过渡树视觉 grounding)解决这一差距,该方法结合了目标状态过渡树(TSTT)与视觉证据 grounding 适配(VEGA)。TSTT将可见交互编译为协议定义的事件,在版本化目标状态上执行这些事件,并将最终查询参考解析到相应的支持证据。VEGA基于轨迹衍生的相同实例对进行适配,使用视觉证据包对解析后的实例执行支持条件 grounding。TT-VG达到70.27%的Joint@0.5指标;在匹配目标解析的情况下,VEGA超出最强对比方法16.20个百分点。在需要路由到非最新或恢复的支持证据的Counter-Recency(反新近性)和Rollback(回滚)任务上,相较于直接推理的增益最大。综上,这些结果确立了SCVIB作为受控测试平台,并强调有效利用解析后的支持证据进行查询侧相同实例定位是多轮个性化定位中的核心挑战。

英文摘要

We introduce editable state-conditioned visual instance binding, a multi-turn localization setting in which several support-defined instances are introduced across turns and protocol-defined state events determine the final target. We instantiate this setting as SCVIB, comprising 1,050 manually verified support--query base pairs and 1,500 episodes spanning five visual domains, three difficulty levels, and four target-state dependency groups. Direct Seq-free inference reaches only 60.13\% Joint@0.5, indicating that resolving the final reference does not ensure effective use of the corresponding visual evidence for query-side localization. We address this gap with TT-VG (Transition-Tree Visual Grounding), which combines a Target-State Transition Tree (TSTT) with Visual Evidence Grounding Adaptation (VEGA). TSTT compiles the visible interaction into protocol-defined events, executes them over versioned target states, and resolves the final-query reference to the corresponding support evidence. Adapted on trajectory-derived same-instance pairs, VEGA performs support-conditioned grounding of the resolved instance using a Visual Evidence Package. TT-VG reaches 70.27\% Joint@0.5; under matched target resolution, VEGA exceeds the strongest comparison method by 16.20 points. Gains over direct inference are largest on Counter-Recency and Rollback, which require routing to non-latest or restored support evidence. Together, these results establish SCVIB as a controlled testbed and highlight the effective use of resolved support evidence for query-side same-instance localization as a central challenge in multi-turn personalized localization.

↑