发表机构
Tsinghua University; Shadow AI(清华大学; 影深AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Tele360提出首个实时前馈系统,从稀疏无位姿相机流中联合估计位姿并重建动态3D高斯,实现2K分辨率下超25 FPS的实时自由视角人体渲染。
AI 中文摘要
实时自由视角可视化真实人类对于沉浸式通信和交互式数字体验至关重要。现有方法要么依赖计算成本高昂的优化,要么需要标定相机和低分辨率输入,使得实时高分辨率部署不切实际。在本工作中,我们提出Tele360,这是首个从稀疏、无位姿RGB流进行动态人体重建和实时自由视角可视化的实时前馈系统。我们的系统在单次前向传播中联合估计相机位姿并为每个时间实例重建动态3D高斯表示。为实现此目标,我们首先设计了一个轻量级的稀疏感知多视图Transformer骨干网络,该网络对前景人体区域进行分词,同时通过共享场景令牌保留全局上下文。然后,我们采用全基于Transformer的高斯解码器,以减轻卷积引起的过度平滑,同时保持解码的稀疏性和高效性。此外,我们引入了一个混合特征金字塔,将多尺度外观线索注入几何预测。我们进一步引入了一个轻量级可微Levenberg-Marquardt相机精化层,以增强多视图一致性和几何对齐。而且,为了在稀疏、无位姿输入下稳定学习,我们通过教师-学生蒸馏从大型视觉几何基础模型迁移多视图几何先验。最后,预测的高斯图通过视频编解码器流式传输到远程设备,用于交互式自由视角渲染。大量实验表明,Tele360在演播室基准上达到了最先进的视觉质量,同时在单个消费级GPU上支持超过25 FPS的实时2K输入到渲染。额外捕获的序列展示了其在我们的多相机设置下对不同主体、服装和动作的性能。
英文摘要
Live free-viewpoint visualization of real humans is critical for immersive communication and interactive digital experiences. Existing methods either rely on computationally expensive optimization or require calibrated cameras and low-resolution inputs, making real-time high-resolution deployment impractical. In this work, we present Tele360, the first real-time feed-forward system for dynamic human reconstruction and live free-viewpoint visualization from sparse, unposed RGB streams. Our system jointly estimates camera poses and reconstructs a dynamic 3D Gaussian representation for each time instance in a single forward pass. To achieve this, we start by designing a lightweight sparsity-aware multi-view transformer backbone that tokenizes foreground human regions while preserving global context through a shared scene token. We then employ a fully transformer-based Gaussian decoder to mitigate convolution-induced over-smoothing while keeping decoding sparse and efficient. In addition, we introduce a hybrid feature pyramid that injects multi-scale appearance cues into geometry prediction. We further introduce a lightweight differentiable Levenberg-Marquardt camera refinement layer to enhance multi-view consistency and geometric alignment. Moreover, to stabilize learning under sparse, unposed inputs, we transfer multi-view geometry priors from a large visual-geometry foundation model via teacher-student distillation. Finally, the predicted Gaussian maps are streamed with video codecs to remote devices for interactive free-viewpoint rendering. Extensive experiments show that Tele360 achieves state-of-the-art visual quality on studio benchmarks while supporting real-time 2K input-to-rendering at over 25 FPS on a single consumer GPU. Additional captured sequences illustrate its performance across varied subjects, clothing, and motions under our multi-camera setup.