arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Struct-GStream:基于结构化3D高斯的低码率高效自由视点视频流技术

Struct-GStream: Towards Efficient Free-Viewpoint Video Streaming at Low-Bitrates with Structured 3D Gaussians

Han Jiao, Jiakai Sun, Lei Zhao, Wei Xing, Huaizhong Lin, Zhanjie Zhang, Ao Ma

arXiv 2608.01053首次发表:更新:

发表机构

College of Computer Science and Technology, Zhejiang University; Chinese Academy of Sciences; University of Science and Technology Beijing(浙江大学计算机科学与技术学院; 中国科学院; 北京科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对现有自由视点视频在线训练方法存储与训练时间不足的问题,提出Struct-GStream,通过结构化3D高斯等技术实现低码率下的高效视频流,性能优于现有方法。

AI 中文摘要

从一组带姿态的2D图像构建动态场景的逼真自由视点视频(FVV)是计算机视觉领域既引人入胜又极具挑战性的任务。基于神经渲染的方法在FVV构建中可实现高保真图像质量,但多数这类方法无法实现实时渲染,且通常需要完整视频序列进行训练。尽管已有部分在线训练方法能实时渲染FVV,但它们难以满足下游应用对存储和训练时间的要求。为解决该问题,本文提出Struct-GStream,可利用结构化3D高斯(3DGs)实现高效FVV流。具体而言,引入动态锚点生成结构化3DGs以构建基础场景,并基于物体运动的局部刚性假设对近似场景运动进行建模;此外,引入全局自由3DGs补丁策略,包含自由3DGs的生成、修剪与优化,以补丁和建模缺陷区域及新出现的物体。所提方法可在低码率下实现快速训练,同时保持高渲染质量。大量实验表明,Struct-GStream在训练时间、存储和渲染质量方面显著优于现有FVV构建的在线训练方法,同时保持有竞争力的渲染速度。

英文摘要

Constructing photorealistic Free-Viewpoint Videos (FVVs) of dynamic scenes from a set of posed 2D images has been an intriguing yet challenging task in computer vision. Methods based on neural rendering achieve high-fidelity image quality in FVV construction. However, most of these methods are unable to achieve real-time rendering and often require complete video sequences to train. Despite the existence of some online training methods capable of rendering FVVs in real time, they struggle to meet the requirements for storage and training time for downstream applications. To overcome this problem, we propose Struct-GStream, which can achieve efficient FVV streaming using structured 3D Gaussians (3DGs). Specifically, we introduce dynamic anchor points to generate structured 3DGs to construct basic scenes and model approximate scene movements based on the assumption of local rigidity in object motion. Besides, we introduce a global free 3DGs patching strategy involving free 3DGs' generation, pruning, and optimization to patch and model deficient areas and emerging objects. Our method achieves fast training at low bitrates while maintaining high rendering quality. Extensive experiments demonstrate that Struct-GStream significantly outperforms existing online training methods for FVV construction in terms of training time, storage, and rendering quality while maintaining competitive rendering speed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑