arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.20639cs.CV

MV2GF:基于视觉几何基础模型的多视角行人检测

MV2GF: Multi-view Pedestrian Detection with a Visual Geometric Foundation Model

Taiga Yamane, Satoshi Suzuki, Ryo Masumura, Shota Orihashi, Tomohiro Tanaka, Mana Ihori, Naoki Makishima

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有多视角行人检测方法泛化性差的问题,提出基于视觉几何基础模型的MV2GF,通过融合任务特定特征与通用几何特征、利用3D点图投影像素,提升了未见过相机配置下的泛化能力。

中文摘要 AI 辅助

多视角行人检测(MVPD)旨在从多视角图像生成的鸟瞰图中检测行人。现有MVPD方法采用统一框架,将2D图像特征投影到3D世界空间并聚合为单一特征,虽有效,但存在两大问题导致训练时难以泛化至未见过的相机配置:一是难以在未见过的相机配置中捕捉跨视角的准确视觉几何;二是因图像特征投影方式,使检测模型高度依赖训练中的失真模式。为解决这些问题,本文利用视觉几何基础模型,提出MV2GF。该基础模型在捕捉跨视角视觉几何、预测不同相机配置下的准确3D属性方面表现出强泛化性。MV2GF将任务特定特征与该基础模型提取的通用几何特征融合,即使在未见过的相机配置中也能有效捕捉视觉几何;此外,MV2GF利用该基础模型预测的3D点图,将图像特征中的每个像素投影到合适的3D位置,避免检测模型依赖训练中的失真模式。实验表明,利用视觉几何基础模型对MVPD的有效性,且MV2GF比现有方法泛化能力更强。

英文摘要

Multi-View Pedestrian Detection (MVPD) aims to detect pedestrians in the form of a bird's eye view map from multi-view images. Recent MVPD methods adopt a unified framework that projects 2D image features into a 3D world space and aggregates them into a single feature. Although they are effective, they struggle to generalize to unseen camera configurations during training due to two main issues. First, they are difficult to capture accurate visual geometry across views in unseen camera configurations. Second, they make detection models highly dependent on distortion patterns during training arising from their image feature projection. To address these, we leverage a visual geometric foundation model and propose MV2GF. This foundation model has exhibited strong generalization in capturing visual geometry across views and predicting accurate 3D attributes in diverse camera configurations. MV2GF fuses task-specific features with general-purpose geometric features extracted by the foundation model to effectively capture the visual geometry even in unseen camera configurations. Furthermore, MV2GF projects each pixel in the image features to an appropriate 3D location using 3D pointmaps predicted by the foundation model, preventing the detection model from depending on distortion patterns during training. Our experiments demonstrate the effectiveness of leveraging a visual geometric foundation model for MVPD and that MV2GF generalizes better than existing methods.

发表机构

  • Human Informatics Laboratories, NTT, Inc.(NTT公司人类信息学实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑