arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07726cs.CV

视觉与WiFi融合:基于物理的体积力学特性估计

Vision Meets WiFi: Physics-Grounded Estimation of Volumetric Mechanical Properties

Ali Bahri, Hongliang Li, Soufiane Lamghari, Jie Chuai, Zhitang Chen

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出ViWi框架,以物体为中心融合视觉与WiFi射频证据,实现体积力学特性估计,在多项基准指标上优于现有方法,提升了估计的准确性与物理一致性。

中文摘要 AI 辅助

仅从视觉估计体积力学特性(包括每个体素的杨氏模量、泊松比和密度)本质上存在歧义,因为视觉相似的物体可能具有截然不同的材料组成和物理行为。现有方法会跨体素独立预测这些特性,忽略了真实物体的分段恒定材料结构,导致相同材料的体素产生嘈杂或不一致的估计,且缺乏明确的机制来解决视觉歧义。我们提出ViWi(Vision Meets WiFi),这是一种以物体为中心的体积力学特性估计框架。ViWi使用一组紧凑的材料槽表示每个物体,这些槽聚合具有相同材料身份的体素的证据,并生成连贯的槽级特性预测。为了补充视觉外观,ViWi纳入了通过WiFi频段电磁模拟生成的紧凑射频(RF)描述符,该描述符利用介电常数和电导率生成。RF描述符为材料槽提供了全局组成线索,而这些线索可能无法从图像中获取,同时视觉特征保留了体素级空间定位。在体积力学特性和质量估计基准测试中,ViWi在6项体素级指标中的4项上优于现有技术,而其仅视觉的变体在所有质量估计指标上均有所提升。这些结果表明,将以物体为中心的材料结构与互补的RF证据相结合,能够实现比仅从视觉外观可能获得的更准确、物理上更一致的体积特性估计。

英文摘要

Estimating volumetric mechanical properties, including Young's modulus, Poisson's ratio, and density at each voxel, is intrinsically ambiguous from vision alone, as visually similar objects may have substantially different material compositions and physical behavior. Existing approaches predict these properties independently across voxels, overlooking the piecewise-constant material structure of real objects and producing noisy or inconsistent estimates for voxels that share the same material, while lacking an explicit mechanism to resolve visual ambiguity. We introduce ViWi (Vision Meets WiFi), an object-centric framework for volumetric mechanical-property estimation. ViWi represents each object using a compact set of material slots that aggregate evidence from voxels with a shared material identity and produce coherent slot-level property predictions. To complement visual appearance, ViWi incorporates a compact RF descriptor generated through WiFi-band electromagnetic simulation using permittivity and conductivity. The RF descriptor conditions the material slots with global composition cues that may be unavailable from images, while visual features preserve voxel-level spatial localization. On GVM, ViWi improves over the prior state of the art on four of six per-voxel metrics, while its vision-only variant improves all reported mass-estimation metrics on ABO-500. These results demonstrate that combining object-centric material structure with complementary RF evidence enables more accurate and physically coherent volumetric property estimation beyond what is possible from visual appearance alone.

↑