arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReferTrack:用于具身视觉跟踪的先指认后跟踪

ReferTrack: Referring Then Tracking for Embodied Visual Tracking

Hanjing Ye, Tianle Zeng, Jiazhao Zhang, Shaoan Wang, Zibo Zhang, Weisi Situ, Yuchen Zhou, Yonggen Ling, Hong Zhang

arXiv 2607.20061首次发表:更新:

发表机构

SUSTech; Tencent Robotics X; Peking University; Futian Laboratory(南方科技大学; 腾讯机器人X实验室; 北京大学; 福田实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对具身视觉跟踪问题,提出ReferTrack先指认后跟踪范式,用单个摄像头,先选目标再解码跟踪航点,通过滑动窗口队列保留运动线索,经数据集协同训练增强识别,在EVT - Bench上达先进单视图性能并验证了迁移能力。

AI 中文摘要

具身视觉跟踪(EVT)要求移动智能体仅使用车载视觉,持续跟踪自然语言描述的特定目标。近期的视觉-语言-动作(VLA)策略虽统一了目标识别和轨迹规划,但其思维链(CoT)推理常在难以监督且与显式图像空间检测弱对齐的抽象空间潜在层中运行。为解决此问题,我们引入ReferTrack,一种使用单个前置摄像头进行具身视觉跟踪的先指认后跟踪范式。我们的模型首先从一组索引边界框中选择目标,然后根据此基于图像的决策解码跟踪航点。为随时间保留目标运动线索,ReferTrack维护先前选定边界框的滑动窗口队列,通过时间-视点-边界框指示器(TVBI)令牌将其几何特征注入视觉历史。我们还通过在自定义Refer-QA数据集上进行协同训练来增强目标识别。在EVT-Bench上,ReferTrack在单目标、分心和模糊跟踪分割上分别实现了89.4%、73.3%和74.1%的成功率,达到了单视图的先进性能,在识别繁重任务上匹配甚至超越了多个多摄像头基线。最后,在有腿和人形机器人上的实际部署验证了其强大的从模拟到现实的迁移能力。

英文摘要

Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) policies unify target identification and trajectory planning, their chain-of-thought (CoT) reasoning often operates in abstract spatial latents that are difficult to supervise and weakly aligned with explicit image-space detections. To address this, we introduce ReferTrack, a referring-then-tracking paradigm that grounds EVT using a single forward-facing camera. Our model first selects the target from an indexed set of bounding boxes, then decodes tracking waypoints conditioned on this image-grounded decision. To preserve target motion cues over time, ReferTrack maintains a sliding-window queue of previously selected bounding boxes, injecting their geometric features into the visual history via temporal-viewpoint-bbox indicator (TVBI) tokens. We further enhance target identification by co-training on a custom Refer-QA dataset. On EVT-Bench, ReferTrack achieves state-of-the-art single-view performance with success rates of 89.4%, 73.3%, and 74.1% on the single-target, distracted, and ambiguity tracking splits, respectively -- matching or even surpassing several multi-camera baselines on identification-heavy tasks. Finally, real-world deployments on legged and humanoid robots validate its robust sim-to-real transfer capabilities. Code is available at https://github.com/MedlarTea/referTrack.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑