arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08826cs.CV

PanoPed:超越边界框的全景行人跟踪的仿真到现实迁移

PanoPed: Beyond Bounding Boxes for Sim-to-Real Panoramic Pedestrian Tracking

Qinfeng Zhu, Weiguang Zhao, Yunxi Jiang, Anh Nguyen, Lei Fan

首次发表
浏览论文内容

中文总结 AI 辅助

针对全景行人跟踪中平面边界框无法准确表示球面位置的问题,提出PanoPed基准和Sextant角度定位头,以极低参数提升跟踪精度,并在仿真到现实迁移中表现优异。

中文摘要 AI 辅助

全景相机使固定监控系统和移动机器人能够在所有方向上跟踪人员,但平面边界框并不能完全描述人在球面上的位置。我们提出了PanoPed,一个用于全球面行人跟踪的仿真到现实基准。PanoPed-S包含来自固定、四足机器人和无人机搭载相机的108,000帧图像,并带有同步的掩码、深度、相机位姿和3D行人状态。PanoPed-R增加了来自固定相机的28,002个真实帧,其中16,247帧被密集标注。我们发现ERP矩形无法唯一确定可见人的球面中心和角范围,而检测器的视觉查询仍携带有关这些信息。受六分仪使用角度测量定位物体的启发,我们提出了Sextant,一个即插即用的角度定位头,仅约0.035M参数。它重用冻结的检测器,保持跟踪身份不变,且不需要额外的图像编码器。Sextant在我们的PanoPed-S测试比较中取得了最佳结果,将最强基线MOTIP从47.30提高到49.49 HOTA,并在所有八个测试序列上均有提升。无需在真实数据上微调,相同的合成训练头在真实视频上将MOTIP和HAT提高了0.96-1.14 HOTA,并且两个种子都改善了每个真实序列。HAT+Sextant在未添加定位图像编码器的比较系统中得分最高。

英文摘要

Full-sphere panoramic cameras let fixed monitoring systems and mobile robots track people in every direction, but a planar bounding box does not fully describe where a person is on the sphere. We introduce PanoPed, a sim-to-real benchmark for pedestrian tracking on the full sphere. PanoPed-S contains 108,000 frames from fixed, quadruped-mounted, and drone-mounted cameras, with synchronized masks, depth, camera poses, and 3D pedestrian states. PanoPed-R adds 28,002 real frames from fixed cameras, 16,247 of them densely annotated. We find that an ERP rectangle cannot uniquely determine the spherical center and angular extent of the visible person, while the detector's visual query still carries information about them. Inspired by the sextant's use of angular measurements to locate objects, we propose Sextant, a plug-and-play angular localization head with only about 0.035M parameters. It reuses a frozen detector, keeps track identities unchanged, and needs no extra image encoder. Sextant gives the best result in our PanoPed-S test comparison, raising the strongest baseline, MOTIP, from 47.30 to 49.49 HOTA, with gains on all eight test sequences. Without fine-tuning on real data, the same synthetic-trained heads improve MOTIP and HAT by 0.96-1.14 HOTA on real video, and both seeds improve every real sequence. HAT+Sextant scores best among the compared systems that add no localization image encoder.

发表机构

  • Xi’an Jiaotong-Liverpool University(西交利物浦大学)
  • University of Liverpool(利物浦大学)
  • Duke Kunshan University(昆山杜克大学)
  • CNRS(法国国家科学研究中心)

机构由 AI 辅助整理,请以论文原文为准。

↑