arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自监督多视角3D注视目标估计:基于概率射线行进

Self-Supervised Multi-View 3D Gaze Target Estimation via Probabilistic Ray Marching

Keqi Chen, Vinkle Srivastav, Nicolas Padoy

arXiv 2609.07415首次发表:更新:

发表机构

University of Strasbourg; IHU Strasbourg; Indian Institute of Technology (IIT) Madras(斯特拉斯堡大学; 斯特拉斯堡大学医院研究所; 印度理工学院马德拉斯分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出自监督方法Self-MVGTE,首次直接估计3D注视目标,利用概率射线行进框架处理噪声伪标签,在MVGT数据集上超越全监督方法。

AI 中文摘要

我们提出了一种自监督方法Self-MVGTE,用于从多个相机视角估计3D注视目标。与现有方法独立估计每个相机视角的2D注视目标不同,Self-MVGTE首次直接在3D空间中预测注视目标。此外,它不需要目标场景的任何真实标注,仅使用来自标定相机设置的多视角输入图像、来自单目注视目标估计模型的伪2D注视目标标签以及来自单目3D注视估计模型的3D注视向量。一个关键挑战是这些伪标签固有地带有噪声且多视角不一致。为解决此问题,我们提出了一种概率射线行进框架,该框架对这些伪标签的不确定性进行建模,并利用3D注视向量作为几何先验。具体而言,这些注视向量首先被集成到单目注视目标估计模型中,以提高其对未见场景的泛化能力,从而产生更高质量的伪标签。然后,对于3D注视目标估计,我们通过从眼睛位置围绕注视向量发射一束射线来构建3D注视锥,以严格约束解空间。在该锥内,我们提出了一种使用现成的DINOv2和Depth-Anything-3模型的深度引导特征采样策略,并估计注视目标的空间似然分布。最后,我们将伪注视目标标签转换为目标分布,并软优化网络。在MVGT数据集上的大量实验表明,Self-MVGTE达到了最先进的性能,超越了现有的全监督基线方法。

英文摘要

We present a self-supervised approach, Self-MVGTE, for estimating 3D gaze targets from multiple camera views. Unlike existing methods that independently estimate 2D gaze targets per camera view, Self-MVGTE predicts gaze targets directly in 3D space for the first time. Moreover, it does not require any ground-truth annotations from the target scene and uses only the multi-view input images from a calibrated camera setup, pseudo 2D gaze target labels from a monocular gaze target estimation model, and 3D gaze vectors from a monocular 3D gaze estimation model. A key challenge is that these pseudo labels are inherently noisy and multi-view inconsistent. To address this, we propose a probabilistic ray marching framework, which models the uncertainty of these pseudo labels and exploits 3D gaze vectors as geometric priors. Specifically, these gaze vectors are first integrated into the monocular gaze target estimation model to improve its generalization to unseen scenes, producing higher-quality pseudo labels. Then, for 3D gaze target estimation, we construct a 3D gaze cone by casting a bundle of rays from the eye position around the gaze vector to strictly constrain the solution space. Within this cone, we propose a depth-guided feature sampling strategy using off-the-shelf DINOv2 and Depth-Anything-3 models, and estimate a spatial likelihood distribution of the gaze target. Finally, we convert the pseudo gaze target labels into a target distribution and softly optimize the network. Extensive experiments on the MVGT dataset show that Self-MVGTE achieves state-of-the-art performance, surpassing existing fully-supervised baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑