HandFlow:基于流匹配的全生成式4D手部恢复
HandFlow: Fully Generative 4D Hand Recovery with Flow Matching
浏览论文内容
中文总结 AI 辅助
针对单目4D手部重建难题,HandFlow提出全生成式流匹配框架,利用ODE积分、双流变压器和掩码机制,在世界空间精度、时间平滑性及重建速度上表现优异,性能达最优。
中文摘要 AI 辅助
准确的单目4D手部重建仍然具有挑战性。逐帧判别回归器缺乏时间上下文,预测不稳定。时间模型虽能聚合信息提高一致性,但易受遮挡和运动模糊影响。生成建模可学习手部运动序列先验,在视觉证据不完整或不可靠时实现连贯手部状态恢复。基于此,提出HandFlow,通过单个ODE积分对MANO参数的整个时间窗口去噪。使用双流变压器捕捉长程依赖,置信感知连续掩码机制处理噪声或缺失观测。实验表明HandFlow性能达最优,在世界空间精度和时间平滑性上提升显著,单GPU上以47fps重建150帧序列,比之前最快方法快约12倍。
英文摘要
Accurate monocular 4D hand reconstruction remains challenging. Per-frame discriminative regressors lack temporal context and often produce jittery predictions. Temporal models improve consistency by aggregating information across frames, but they are typically deterministic regressors, making them vulnerable to ambiguous observations caused by occlusion and motion blur. Generative modeling offers a natural alternative by learning a prior over plausible hand motion sequences, enabling coherent hand-state recovery when visual evidence is incomplete or unreliable. Motivated by this observation, we present HandFlow, a fully generative flow-matching framework for temporally coherent 3D hand pose and shape estimation from monocular video. Given visual and skeletal observations, HandFlow denoises an entire temporal window of MANO parameters through a single ODE integration. To support this, we use a Flux-style dual-stream transformer that attends across the full sequence to capture long-range dependencies without autoregressive decoding, and a confidence-aware continuous masking mechanism that blends observed features with learnable mask tokens to handle noisy or missing observations. Experiments on DexYCB and HOT3D show that HandFlow achieves state-of-the-art performance, with particularly large gains in world-space accuracy and temporal smoothness. It reduces world-space pose error by over 30% compared with the strongest baseline and achieves the lowest acceleration error among all evaluated methods, while remaining competitive in per-frame pose accuracy. Moreover, on a single GPU HandFlow reconstructs a 150-frame sequence at 47 fps, about 12x faster than the fastest prior video-based method, with reconstruction itself accounting for only a small fraction of the end-to-end latency.
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Google(谷歌)
机构由 AI 辅助整理,请以论文原文为准。