发表机构
Michigan State University(密歇根州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究从单图像推断物理属性,核心方法是引入SiPhy框架,将视觉线索与材料知识对齐,通过采样、提取特征等操作及对比聚合器、重量感知细化来实现。在多个数据集上取得领先性能,还验证了其在真实交互数据集上作为数据标注引擎的潜力。
AI 中文摘要
从单张图像推断质量、刚度和弹性等物理属性对模拟和具身人工智能至关重要,但现有方法大多依赖多视图重建或基于物理的监督。我们引入SiPhy,一个用于单图像物理属性推理的统一框架,它将3D感知视觉线索、深度与基于语言的材料知识对齐。从一张RGB图像中,SiPhy对伪体素点进行采样,提取CLIP特征,并将其与VLM提出的材料候选进行关联。基于部分的对比聚合器增强区域一致性,而重量感知细化改进密集物体的厚度和体积估计。在ABO - 500、MVImgNet - 100和PhysXNet - 100数据集上,SiPhy取得了领先的单图像性能,在质量MnRE上比PUGS提高了93%,在密度MAE上比NeRF2Physics降低了35.5%,在杨氏模量误差上降低了23.5%。我们还在真实的手 - 对象交互数据集上验证了SiPhy,展示了其作为从单视图图像进行物理理解的数据标注引擎的潜力。
英文摘要
Inferring physical properties such as mass, stiffness, and elasticity from a single image is essential for simulation and embodied AI, yet most existing approaches rely on multi-view reconstruction or physics-based supervision. We introduce SiPhy, a unified framework for single-image physical property reasoning that aligns 3D-aware visual cues, depth with language-based material knowledge. From one RGB image, SiPhy samples pseudo-voxel points, extracts CLIP features, and grounds them to material candidates proposed by a VLM. A part-based contrastive aggregator enforces region consistency, while a heaviness-aware refinement improves thickness and volume estimation for dense objects. Across ABO-500, MVImgNet-100, and PhysXNet-100, SiPhy achieves state-of-the-art single-image performance, surpassing multi-view reconstruction methods by improving mass MnRE by up to 93% (vs. PUGS), reducing density MAE by 35.5% (vs. NeRF2Physics), and lowering Young's modulus error by 23.5%. We further validate SiPhy on real hand-object interaction datasets, demonstrating its potential as a data annotation engine for physical understanding from single-view imagery.
CommentsAccepted to ECCV 2026 (main track)