发表机构
Nankai University; Beihang University(南开大学; 北京航空航天大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对3D高斯泼溅新视角合成,提出视图结构化共形预测(VSCP),分解尺度为空间形状与视图难度因子,通过视图级校准保证覆盖率并削减预测宽度,实验验证其有效性与高效性。
AI 中文摘要
3D高斯泼溅(3DGS)能够实时渲染新视角,但不确定性热图并不能保证渲染视图达到特定的预测覆盖率。我们将新视角合成视为结构化回归问题,并要求以至少$1-\alpha$的概率,RGB预测框覆盖新视图中至少$1-\beta$比例的像素。我们提出视图结构化共形预测(VSCP)。该方法将预校准尺度分解为来自渲染器的空间形状和可迁移的视图难度因子,该因子预测形状所需的最小视图级乘数。随后,对视图进行留出分位数(View-CP)校准,即使迁移到新场景也能保证有限样本有效性。相同的分解使得分析精确:一致性分数是真实视图难度与预测视图难度之比,多余宽度分为测试端项和校准端项。在13个真实场景中,像素池化校准在90%目标下达到89.9%的边缘像素覆盖率,但视图事件覆盖率仅为61.4%,而View-CP达到91.7%至92.0%。在匹配覆盖率下,VSCP相比恒定尺度将宽度削减22.1%,并且仅使用每场景一个模型和每次查询四次(而非十次)光栅化遍历,即可匹配十个模型集成的21.0%削减效果。VSCP还比最接近的单模型基线3DGS-U场提升了4.7个百分点($p=0.0225$)。视图预测器可从有界源族迁移到所有九个无界Mip-NeRF 360场景。在这些场景中,完整尺度在所有九个场景上比恒定尺度节省20.7%的宽度。在不同的致密化骨干下,它仍保持18.3%的节省,并在RTX 4090上以216至280 FPS运行。
英文摘要
3D Gaussian Splatting (3DGS) renders novel views in real time, but an uncertainty heatmap does not certify that a rendered view meets a certain prediction coverage. We treat novel-view synthesis as structured regression and ask that, with probability at least $1-α$, RGB prediction boxes cover at least a $1-β$ fraction of pixels in a new view. We propose View-Structured Conformal Prediction (VSCP). It splits the pre-calibration scale into a spatial shape from the renderer and a transferable view-difficulty factor, which predicts the smallest view-wise multiplier that shape needs. A held-out quantile over views (View-CP) then gives finite-sample validity even when transferring to new scenes. The same factorization makes the analysis exact: a conformity score is the ratio of oracle to predicted view difficulty, and excess width separates into a test-side and a calibration-side term. Across 13 real scenes, pixel-pooled calibration reaches 89.9\% marginal pixel coverage but only 61.4\% view-event coverage at a 90\% target, while View-CP reaches 91.7--92.0\%. At matched coverage VSCP cuts width by 22.1\% against a constant scale, and matches a ten-model ensemble's 21.0\% reduction using only one model per scene and four rather than ten rasterization passes per query. VSCP also improves on the closest single-model baseline, the 3DGS-U field, by 4.7 points ($p=0.0225$). The view predictor transfers from bounded source families to all nine unbounded Mip-NeRF~360 scenes. There the full scale beats the constant scale with 20.7\% width saving on all nine scenes. It also keeps an 18.3\% saving under a different densification backbone and runs at 216--280 FPS on an RTX~4090.