arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TOM-GS:通过静态3D高斯的时间不透明度调制实现可编辑视频表示

TOM-GS: Editable Video Representation via Temporal Opacity Modulation of Static 3D Gaussians

Marek Lisowski, Łukasz Smoliński, Kornel Howil, Piotr Biliński, Marcin Mazur, Przemysław Spurek

arXiv 2607.22717首次发表:更新:

发表机构

University of Warsaw; Jagiellonian University; IDEAS Research Institute(华沙大学; 雅盖隆大学; IDEAS 研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视频表示难以编辑的问题,提出TOM-GS方法,通过为3D高斯的不透明度赋予可学习参数,结合姿态估计保持静态空间几何,实现可编辑视频表示,在视觉保真度上超越同类,且与3D编辑工具兼容。

AI 中文摘要

虽然隐式神经表示(INRs)和动态3D高斯渲染(3DGS)在视频处理中取得了令人瞩目的成果,但它们生成的表示往往难以编辑。近期方法通过引入复杂空间变形或折叠分布来解决此问题,这限制了优化并降低了下游编辑的灵活性。本文介绍了TOM-GS,一种可编辑视频表示方法,它舍弃复杂变形,采用配备连续时间不透明度公式的常规3D高斯。通过为每个高斯的不透明度分配可学习的时间均值和比例,模型使静态3D空间组件能在场景中平滑淡入淡出。基于强大的现成姿态估计,该方法保持静态空间几何结构,自然支持各种手动和基于物理的编辑。TOM-GS在视觉保真度上优于先前的可编辑视频表示,且因其依赖标准3D高斯,确保与现有3D编辑工具无缝兼容。

英文摘要

While Implicit Neural Representations (INRs) and dynamic 3D Gaussian Splatting (3DGS) achieve impressive results in video processing, they often fall short of producing representations that are easily editable. Recent methods address this by introducing complex spatial deformations or folded distributions, which constrain optimization and reduce flexibility for downstream editing. In this paper, we introduce TOM-GS, an editable video representation that forgoes complex deformations in favor of regular 3D Gaussians equipped with a continuous temporal opacity formulation. By assigning a learnable temporal mean and scale to the opacity of each Gaussian, our model enables static 3D spatial components to fade smoothly in and out of the scene. Grounded by robust, off-the-shelf pose estimation, our approach maintains a static spatial geometry that naturally supports a wide range of manual and physics-based edits. TOM-GS outperforms prior editable video representations in visual fidelity, while its reliance on standard 3D Gaussians ensures seamless compatibility with established 3D editing tools.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑