arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08782cs.CVcs.AIcs.GR

4D-HOF:用于前馈4D交互重建的手-物体流匹配

4D-HOF: Hand-Object Flow Matching for Feed-Forward 4D Interaction Reconstruction

  • NVIDIA(英伟达)
  • Texas A&M University(德克萨斯A&M大学)

机构由 AI 辅助整理,请以论文原文为准。

Shiqi Li, Sean Cho, Yijie Li, Fengzhi Guo, Bowen Wen, Cheng Zhang

AI总结:

4D-HOF提出一种前馈流匹配框架,利用视觉基础模型的粗略估计重建4D手-物体交互,通过条件流匹配和测试时引导实现稳定准确的生成式重建。

AI中文摘要:

现有的4D手-物体重建方法通常依赖于昂贵的逐序列优化,而生成式方法则通常从随机噪声中合成交互,这可能导致不稳定的交互预测。我们提出4D-HOF,一种前馈框架,从视觉基础模型产生的粗略但信息丰富的估计中重建4D手-物体交互。具体来说,我们学习一个条件流匹配模型,将基础模型导出的手-物体状态输送到交互流形上,使模型能够以前馈方式纠正平移、旋转和对齐中的错误。我们生成式公式的一个关键优势是,它自然地在传输过程中实现测试时引导。我们不是在进行重建后应用单独的后处理优化,而是直接使用物理交互约束和观察到的2D证据来引导演化的生成状态,使重建能够在生成过程本身中作为一部分进行细化。通过在多样化数据集上训练生成模型,4D-HOF对具有挑战性的野外场景具有鲁棒泛化能力。在域外基准上的实验表明,4D-HOF达到了最先进的性能,产生更稳定和准确的4D手-物体重建。

英文摘要:

Existing methods for 4D hand-object reconstruction often rely on costly per-sequence optimization, while generative approaches typically synthesize interactions from random noise, which can lead to unstable interaction prediction. We introduce 4D-HOF, a feed-forward framework that reconstructs 4D hand-object interactions from coarse but informative estimates produced by vision foundation models. Concretely, we learn a conditional flow matching model that transports foundation-model-derived hand-object states toward an interaction manifold, allowing the model to correct errors in translation, rotation, and alignment in a feed-forward manner. A key advantage of our generative formulation is that it naturally enables test-time guidance within the transport process. Rather than applying a separate post-hoc optimization after reconstruction, we directly steer the evolving generative states using physical interaction constraints and observed 2D evidence, allowing the reconstruction to be refined as part of the generative process itself. By training the generative model on diverse datasets, 4D-HOF generalizes robustly to challenging in-the-wild scenarios. Experiments on out-of-domain benchmarks show that 4D-HOF achieves state-of-the-art performance, producing more stable and accurate 4D hand-object reconstructions.

补充信息

↑