HyperGS:快速且通用的高斯视频表示
HyperGS: Fast and Generalizable Gaussian Video Representation
浏览论文内容
中文总结 AI 辅助
研究针对高斯视频表示中现有方法依赖单视频优化、编码慢且通用性受限的问题,提出HyperGS方法,通过设计特定Transformer提取令牌和获得高斯表示,结合正则化器解决退化问题,实现快速编码及跨视频通用,提升了PSNR。
中文摘要 AI 辅助
高斯渲染已成为一种有效的视频表示方法,但现有方法依赖于每个视频的优化。这导致编码速度慢并限制了跨视频的通用性。为了分摊这种优化,我们提出了HyperGS,一种前馈、无需优化的方法,它可以在一次前向传递中直接从任何视频预测高斯表示,将编码和解码速度提高几个数量级,同时以更高分辨率推广到分布外视频。在HyperGS中,我们设计了一个因式分解的时空Transformer从视频中提取令牌,并设计了一个基于可学习查询的Transformer为每个视频帧获得8参数高斯表示。我们发现,在不同视频上天真地预测高斯会导致针状退化,从而使训练崩溃,我们用基于秩的几何正则化器来解决这个问题,其强度动态调整以稳定优化。HyperGS在匹配重建质量的情况下,编码速度比每个视频的高斯优化快10^4-10^5倍,同时零样本推广到720p视频,无需重新编码即可实现更高分辨率的渲染。在较小的视频表示尺寸下,HyperGS在K400、SSv2和UCF101上比之前的视频编码器提高了2.9-3.1dB的PSNR。通过在一次前向传递中预测显式的2D高斯,HyperGS将高斯渲染的快速、灵活渲染与前馈预测的速度和通用性结合起来,推动高斯成为快速且通用的视频表示的实用方向。
英文摘要
Gaussian Splatting has emerged as an effective representation for video, but existing methods rely on per-video optimization. This leads to slow encoding and limits generalization across videos. To amortize this optimization, we propose HyperGS, a feedforward, optimization-free approach that directly predicts Gaussian representations from any video in a single forward pass, speeding up encoding and decoding by orders of magnitude while generalizing to out-of-distribution videos at higher resolutions. In HyperGS, we design a factorized spatiotemporal Transformer to extract tokens from video, and a learnable query-based Transformer to obtain 8-parameter Gaussian representations for each video frame. We find that naively predicting Gaussians across diverse videos induces a needle-like degeneration that collapses training, and address this with a rank-based geometric regularizer whose strength adapts dynamically to stabilize optimization. HyperGS achieves encoding at $10^4$--$10^5\times$ the speed of per-video Gaussian optimization at matched reconstruction quality while generalizing zero-shot to $720p$ video, enabling higher-resolution rendering without re-encoding. HyperGS improves PSNR by +2.9--3.1 dB over the prior video encoders on K400, SSv2, and UCF101 at a smaller video representation size. By predicting explicit 2D Gaussians in a single forward pass, HyperGS combines the fast, flexible rendering of Gaussian Splatting with the speed and generalization of feedforward prediction, advancing Gaussians as a practical direction for fast and generalizable video representation.
发表机构
- King Abdullah University of Science and Technology (KAUST)(阿卜杜拉国王科技大学)
- Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。