arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11096cs.CVcs.AI

水陆两栖机器人跟踪的连续真实值构建与恢复策略

Continuous Ground-Truth Construction and a Recovery Policy for Air--Water Robotic Tracking

Jiangong Xiao, Zhe Sun, Kanzhong Yao, Yuanbo Bi, Haofei Zhao, Ruixuan Hu, Guan Huang, Xuelong Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对气水界面视觉跟踪的观测失效问题,本文构建了22346帧的跨介质测试集,并提出Cross-Medium Recovery Policy,使MixFormerV2的跟踪性能显著提升。

中文摘要 AI 辅助

气水界面的视觉跟踪面临水花、气泡、折射、反射及外观突变的挑战,这些问题会暂时使观测失效。该场景带来两个耦合难题:其一,用于评估时,仅基于图像的标注无法可靠描述视觉盲区中目标的物理位置;其二,用于在线跟踪时,受损观测会污染运动估计和外观模板。我们通过一套构建流程解决第一个难题:将相机帧与动作捕捉位姿同步,投影已知目标几何,用介质门控残差校正水下投影,并对标注进行人工审核,由此得到一个包含22346帧的仅用于评估的跨介质测试集。我们进一步提出以置信度触发模板选择为核心的跨介质恢复策略(Cross-Medium Recovery Policy, CMRP),它为MixFormerV2提供固定初始模板、窗口最优预触发模板及触发帧卡尔曼引导的图像裁剪,以及对应的权重,无需对视觉骨干网络进行重新训练。在准确率评估中,CMRP取得49.90的宏成功率AUC,比官方版MixFormerV2高2.95个百分点;在选定的跨介质过渡和遮挡恢复区间,相较于官方更新,CMRP将MixFormerV2的跟踪覆盖率从47.91%提升至50.43%,同时成功恢复视频的平均损失至恢复延迟从55.3帧降至49.3帧。

英文摘要

Visual tracking across the air-water interface is challenged by splashes, bubbles, refraction, reflections, and abrupt appearance changes that can temporarily invalidate observations. This setting poses two coupled difficulties: first, for evaluation, image-only annotation cannot reliably describe the target's physical location during visual blindness; second, for online tracking, corrupted observations can contaminate motion estimates and appearance templates. We address the first difficulty with a construction pipeline that synchronizes camera frames with motion-capture poses, projects known target geometry, corrects underwater projection with a medium-gated residual, and subjects the annotations to manual review. This yields an evaluation-only cross-medium test set of 22,346 frames. We further introduce a Cross-Medium Recovery Policy (CMRP) centered on confidence-triggered template selection. It supplies MixFormerV2 with the fixed initial template, a window-best pre-trigger template, and a trigger-frame Kalman-guided image crop, together with their associated weights, without retraining the visual backbone. In the accuracy evaluation, CMRP achieves 49.90 Macro Success AUC, 2.95 points above MixFormerV2 Official. On selected cross-medium transition and occlusion-recovery intervals, CMRP increases MixFormerV2 tracking coverage from 47.91\% to 50.43\% relative to Official updating, while mean loss-to-recovery latency over successfully recovered videos decreases from 55.3 to 49.3 frames.

发表机构

  • Northwestern Polytechnical University(西北工业大学)
  • Institute of Artificial Intelligence (TeleAI), China Telecom(中国电信人工智能研究院(TeleAI))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑