发表机构
Guangdong Laboratory of Artificial Intelligence and Digital Economy (SZ); Shenzhen University(广东省人工智能与数字经济实验室(深圳); 深圳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MATE4D是一种基于单图像的四维内容生成框架,通过构建时空多视图图像矩阵优化三维高斯基元并结合轻量形变模块,在多数据集上优于基线,可实现高质量可编辑动态4D内容生成,支持AR/VR创作。
AI 中文摘要
生成模型已快速将内容创作从二维图像拓展至动态三维与四维场景合成,但从单张图像生成逼真且时序稳定的四维内容仍具挑战,因为单视图仅能提供有限的结构线索与微弱的运动证据。本文提出MATE4D,一个可将单张输入图像转换为可编辑动态四维内容的框架。该方法通过文本引导的背景操作构建时空多视图图像矩阵,为视角、外观与运动提供一致的监督;这些合成观测结果被用于优化三维高斯基元,随后通过轻量形变模块进行动画处理,形成四维表示。生成的场景能更忠实地保留几何结构,维持更平滑的时序行为,且背景编辑更一致,减少了上下文歧义与运动伪影。在Objaverse-XL与Diffusion4D上的实验表明,MATE4D在视觉质量、效率与可控性方面均优于强基线方法,可支持实用的AR/VR内容创作。
英文摘要
Generative models have rapidly pushed content creation be-yond 2D imagery toward dynamic 3D and 4D scene synthesis. Yet pro-ducing realistic and temporally stable 4D content from a single image is still difficult because one view provides limited structural cues and weak motion evidence. We introduce MATE4D, a framework that converts one input image into editable dynamic 4D content. Our method constructs a spatio-temporal multi-view image matrix with text-guided background manipulation, delivering coherent supervision over viewpoint, appear-ance, and motion. These synthesized observations are used to optimize 3D Gaussian primitives, which are then animated through a lightweight deformation module to form a 4D representation. The resulting scenes preserve geometry more faithfully, maintain smoother temporal behavior, and keep background edits more consistent, reducing context ambiguity and motion artifacts. Experiments on Objaverse-XL and Diffusion4D show that MATE4D outperforms strong baselines in visual quality, effi-ciency, and controllability, supporting practical AR/VR content creation.