基于卷帘快门视频的扫描线感知可动画高斯化身
Scanline-Aware Animatable Gaussian Avatars from Rolling-Shutter Videos
- University of Macau(澳门大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出RS-Avatar,直接从卷帘快门视频重建无失真可动画3D高斯化身,在自建基准数据集RS-ZJU上的新视角合成任务中性能优于忽略快门的基线方法。
AI中文摘要:
可动画人类化身通常通过多视图视频重建,其隐含假设是帧的每个像素都观测到身体运动的同一时刻。卷帘快门(RS)传感器按顺序曝光图像行,因此在一帧内,运动者的头部和脚部会因数十毫秒的关节运动产生错位,每条扫描线对应不同姿态。将此类视频输入现有化身模型会将失真烘焙到规范表示中,导致在新视角和新姿态下出现剪切和抖动问题。更严重的是,装置中的每个相机遵循各自的读取时序,即使几何结构正确,也会破坏重建所需的多视图一致性。本文提出RS-Avatar,可直接从RS视频重建清晰、无失真的可动画3D高斯化身。该方法的公式极为简洁:运动感知化身已能在多个子帧时刻渲染身体,其中模糊模型对这些渲染结果取平均,而卷帘快门模型则按扫描线逐行组合这些渲染结果,仅需修改该操作符即可实现。在我们从ZJU-MoCap构建的基准数据集RS-ZJU上,与假设帧为瞬时进行训练的方法相比,该方法在所有主体的新视角合成任务上均取得了性能提升;而基于相同子帧机制构建的运动感知模糊模型无法迁移,性能甚至低于忽略快门的基线,表明机制可复用,但操作符不可直接复用。
英文摘要:
Animatable human avatars are routinely reconstructed from multi-view video under a silent assumption: that every pixel of a frame observes the same instant of the body's motion. Rolling-shutter (RS) sensors expose image rows sequentially, so within one frame the head and the feet of a moving person are separated by tens of milliseconds of articulated motion, and every scanline sees a different pose. Feeding such video to a state-of-the-art avatar bakes the distortion into the canonical representation, where it survives as shear and wobble under novel views and novel poses. Worse, every camera in a rig follows its own readout schedule, so the multi-view consistency that drives the reconstruction is violated even when the geometry is correct. We present RS-Avatar, which reconstructs a sharp, undistorted, animatable 3D Gaussian avatar directly from RS video. The formulation is minimal: a motion-aware avatar already renders the body at several sub-frame instants, and where a blur model averages those renderings, a rolling-shutter model composites them scanline by scanline. Changing that operator is the only modification required. On RS-ZJU, a benchmark we build from ZJU-MoCap, this improves novel-view synthesis over training as if the frames were instantaneous, on every subject. A motion-aware blur model built on the same sub-frame machinery does not transfer, and in fact falls below the shutter-oblivious baseline: the machinery is reusable, the operator is not.