PhysVGGT:从单张图像进行前馈式稠密物理属性估计
PhysVGGT: Feed-Forward Dense Physical Property Estimation from A Single Image
- Concordia University(康考迪亚大学)
- Huawei(华为)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
PhysVGGT提出一种前馈式稠密物理属性估计方法,从单张RGB图像预测摩擦、硬度等属性图及物体质量,无需逐物体重建,推理速度快27倍,在ABO-500和NeRF2Physics数据集上达到最先进性能。
AI中文摘要:
物理属性,如摩擦力、硬度、刚度和密度,决定了机器人应如何抓取、操控和与物体交互,然而从RGB图像中估计这些属性仍然具有挑战性。现有方法通常采用逐物体重建并结合物理属性,或在测试时直接查询视觉语言模型,这导致大量计算开销,限制了其适用性。在这项工作中,我们提出了PhysVGGT,一种前馈模型,能够在一次前向传播中从单张RGB图像预测摩擦系数、邵氏硬度、杨氏模量和密度的稠密图,以及物体级别的质量。PhysVGGT的关键思想是将物理属性估计表述为稠密的逐像素预测问题,并采用视觉几何变换器从输入图像中提取几何感知的标记,随后通过稠密预测分支估计局部物理属性,通过全局预测分支估计物体级别的质量。此外,我们引入了一个可扩展的伪标签生成流程,能够实现大规模弱监督训练,用于稠密物理属性预测,大幅减少了对昂贵的直接物理测量的需求。大量实验表明,PhysVGGT在ABO-500数据集上达到了最先进的性能,并有效泛化到分布外的NeRF2Physics数据集。此外,PhysVGGT无需逐物体重建和测试时优化,每张图像的推理延迟仅为0.13秒,比之前的最先进方法快27倍。
英文摘要:
Physical properties, such as friction, hardness, stiffness, and density, govern how robots should grasp, manipulate and interact with objects, yet estimating these properties from RGB images remains challenging. Existing methods typically employ per-object reconstruction augmented with physical properties or directly query vision-language models at test time, which results in substantial computational overhead that limits their applicability. In this work, we present PhysVGGT, a feed-forward model that predicts dense maps of friction coefficient, Shore hardness, Young's modulus, and density, together with object-level mass, from a single RGB image in one forward pass. The key idea of PhysVGGT is to formulate physical property estimation as a dense per-pixel prediction problem and employ a visual geometry transformer to extract geometry-aware tokens from the input image followed by a dense prediction branch for estimating local physical properties and a global prediction branch for estimating object-level mass. In addition, we introduce a scalable pseudo-label generation pipeline that enables large-scale weakly supervised training for dense physical property prediction, substantially reducing the need for expensive direct physical measurements. Extensive experiments show that PhysVGGT achieves state-of-the-art performance on the ABO-500 dataset and generalizes effectively to the out-of-distribution NeRF2Physics dataset. Moreover, PhysVGGT eliminates the need for per-object reconstruction and test-time optimization, achieving an inference latency of only 0.13s per image, making it $27\times$ faster than the previous state of the art.