发表机构
Moholo Inc.(Moholo公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对3D高斯泼溅中球谐系数内存开销大且视角冗余的问题,提出基于观测Gram矩阵的失真度量,实现阶数降低、分配和向量量化的统一优化,在无需训练下提升压缩率与质量。
AI 中文摘要
3D高斯泼溅模型的大部分内存用于存储球谐颜色系数,然而每个高斯仅从训练相机的狭窄视角锥中被观测到。我们将此转化为一种其他压缩器可采用的失真度量:一个逐高斯的观测Gram矩阵,由视角方向和混合权重累积而成,是系数变化到平方图像误差的精确一阶映射,且仅需模型和相机位姿。在该度量下,阶数降低成为封闭形式的投影,推广了截断;阶数分配成为拉格朗日率失真问题;向量量化成为矩阵加权的Lloyd算法,其中Compressed3D的量化器是其标量情形。将该度量替换进Compressed3D且其他一切不变,在微调前PSNR提升+0.49 dB,SSIM和LPIPS也随之提升,在匹配码率下仍获得+0.32 dB增益,且无需任何训练图像。仅基于该度量的免训练堆栈在Mip-NeRF 360上以相同质量比无图像GSICO小15%。
英文摘要
Most of the memory of a 3D Gaussian Splatting model holds spherical-harmonic colour coefficients, yet each Gaussian is seen only from the narrow cone of directions of the training cameras. We turn this into a distortion metric that other compressors can adopt: a per-Gaussian observation Gram matrix, accumulated from viewing directions and blending weights, is the exact first-order map from coefficient changes to squared image error and needs only the model and the camera poses. Under it, degree reduction becomes a closed-form projection that generalises truncation, degree allocation a Lagrangian rate-distortion problem, and vector quantisation the matrix-weighted Lloyd algorithm, of which Compressed3D's quantiser is the scalar case. Swapped into Compressed3D with everything else unchanged, the metric raises PSNR by +0.49 dB before fine-tuning, with SSIM and LPIPS following, and at matched rate still gains +0.32 dB without a single training image. A training-free stack built on the metric alone is 15% smaller than the image-free GSICO at equal quality on Mip-NeRF 360.
Comments22 pages, 6 figures