arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29106cs.CVcs.AIcs.CG

WildHSR:基于3D基础模型的度量前馈4D人物-场景重建

WildHSR: Metric Feed-Forward 4D People-Scene Reconstruction from a 3D Foundation Model

Jerrin Bright, John Zelek

首次发表
浏览论文内容

中文总结 AI 辅助

提出WildHSR,利用3D基础模型,通过尺度伪标签预训练和轻量适配实现度量尺度预测,并利用中间层特征进行人物关联,实现前馈式4D人物-场景重建,在EMDB-2和RICH上取得领先性能。

中文摘要 AI 辅助

3D基础模型在一次前向传播中恢复视频相机和几何,但一些最强的模型结果存在尺度不确定性。联合人物-场景重建需要两个缺失的输出:度量尺度和持久的人物身份。我们探究一个尺度不确定的基础表示是否可以通过轻量级适配来支持两者。精确的度量标签稀缺,但未标注的真实世界视频丰富。我们利用精选网络视频中的人物来初始化解决方案:一个姿态度量身体和2D关键点提供近似的闭式尺度伪标签。这些伪标签预训练一个尺度读出器,然后使用来自标准真实视频训练划分的精确度量监督,与轻量级适配器一起进行微调。在推理时,该头部从基础模型标记预测度量尺度,无需标尺或其教师。对于人物身份,我们单独探测预训练的基础模型,发现其中间查询-键特征编码了跨帧的人物对应关系。在大多数评估的移动人物片段中,中间层标记优先于该人物而非空置位置和其他人物。一个微小的投影读取此对应关系;结合度量骨盆运动和提议置信度,它驱动了带垃圾箱感知的Sinkhorn关联,用于逐帧身体。WildHSR结合两个读出器,从单目视频重建度量相机、场景和人物。每个窗口都是前馈预测;解析关联和Sim(3)组合连接窗口。在EMDB-2上,WildHSR是已发表比较中第一个在WA-MPJPE和RTE上击败最佳基于优化的方法的前馈方法,同时在所有三个世界坐标系指标上领先前馈方法。在RICH上,它在WA-MPJPE和W-MPJPE上领先前馈人物-场景方法。完整流程在单GPU上以10.1帧/秒运行。

英文摘要

3D foundation models recover video cameras and geometry in one forward pass, but some of the strongest are up to scale. Joint people-scene reconstruction then requires two missing outputs: metric scale and persistent person identity. We ask whether one up-to-scale foundation representation can support both through lightweight adaptation. Exact metric labels are scarce, but unlabeled in-the-wild video is abundant. We use people in curated web video to initialise the solution: a posed metric body and 2D keypoints give an approximate, closed-form scale pseudo-label. These pseudo-labels pretrain a Scale Readout, which is then fine-tuned together with a lightweight adapter using exact metric supervision from standard real-video training splits. At inference the head predicts metric scale from foundation-model tokens, without the ruler or its teachers. For person identity, we probe the pretrained foundation model alone and find evidence that its intermediate query-key features encode person correspondence across frames. In most evaluated moving-person clips, a mid-layer token prefers that person over the vacated location and other people. A tiny projection reads this correspondence; together with metric pelvis motion and proposal confidence, it drives dustbin-aware Sinkhorn association of per-frame bodies. WildHSR combines both readouts to reconstruct metric cameras, scene and people from monocular video. Each window is predicted feed-forward; analytic association and Sim(3) composition connect windows. On EMDB-2, WildHSR is the first feed-forward method in the published comparison to beat the best optimization-based WA-MPJPE and RTE while leading feed-forward methods on all three world-frame metrics. On RICH, it leads feed-forward people-and-scene methods on WA-MPJPE and W-MPJPE. The complete pipeline runs at 10.1 fps on one GPU.

发表机构

  • Vision and Image Processing Lab, University of Waterloo(滑铁卢大学视觉与图像处理实验室)

机构由 AI 辅助整理,请以论文原文为准。

↑