arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22731cs.CV

室内环境下无校准的3D多摄像机人体跟踪

Calibration-Free 3D Multi-Camera People Tracking for Indoor Environment

Ponleur Veng, Dominique Vaufreydaz, Phutphalla Kong

首次发表
浏览论文内容

中文总结 AI 辅助

针对室内环境下多摄像机人体跟踪依赖校准的问题,提出统一无校准3D MCPT框架,集成多种技术,采用姿态引导3D提升策略和全局身份关联方法,在挑战赛中取得53.13%的HOTA分数,为纯视觉3D跟踪建立基线。

中文摘要 AI 辅助

多摄像机人体跟踪(MCPT)传统上依赖精确的相机内参和外参校准,将2D检测投影到统一的3D世界坐标中。手动校准是从无约束视频档案生成大规模数据集的主要瓶颈。本文提出了一个统一的无校准3D MCPT框架,该框架使用深度基础模型直接从视觉数据推断几何结构。系统集成了无锚点检测(YOLOX)、鲁棒跟踪(BoT-SORT)、全尺度外观嵌入(OsNet)、姿态估计(通过MMPose的HRNet)以及使用视觉几何基础变换器(VGGT)的基于变换器的几何重建。一种姿态引导的3D提升策略将头部关键点投影到重建流形上,消除了对地面平面单应性的依赖。全局身份关联被制定为在具有严格速度门控的联合外观-几何成本下的层次凝聚聚类。在2024年人工智能城市挑战赛上的评估表明,在无法获取地面真值校准矩阵的情况下,HOTA分数为53.13%,为基于纯视觉的3D跟踪建立了一个强大的基线。

英文摘要

Multi-Camera People Tracking (MCPT) traditionally relies on precise intrinsic and extrinsic camera calibration to project 2D detections into a unified 3D world coordinate system.However, manual calibration constitutes a major bottleneck in large-scale dataset generation from unconstrained video archives. This work proposes a unified calibration-free 3D MCPT framework that infers geometric structure directly from visual data using deep foundation models. The system integrates anchor-free detection (YOLOX), robust tracking (BoT-SORT), omni-scale appearance embedding (OsNet), pose estimation (HRNet via MMPose), and transformer-based geometric reconstruction using the Visual Geometry Grounded Transformer (VGGT). A pose-guided 3D lifting strategy projects head keypoints onto a reconstructed manifold, eliminating dependence on ground-plane homography. Global identity association is formulated as hierarchical agglomerative clustering under a joint appearance-geometry cost with strict velocity gating. Evaluation on the AI City Challenge 2024 demonstrates a HOTA score of 53.13% without access to ground-truth calibration matrices, establishing a strong baseline for purely vision-based 3D tracking.

补充信息

↑