arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ManiPhysicsBench:VLA操作中物体保持的基于物理学的评估

ManiPhysicsBench: Physics-Based Assessment of Object Preservation in VLA Manipulation

Sangwu Park, Yeonjun In, Wonjoong Kim, Sungwon Kim, Sein Kim, Chanyoung Park

arXiv 2610.02802首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出ManiPhysicsBench,基于物理模拟评估VLA模型在操作中能否保持物体,发现公开模型安全成功率低,重训练后有所提升但泛化受限。

AI 中文摘要

视觉-语言-动作(VLA)模型旨在执行多样化的操作任务,但现有刚体基准中的任务成功并不表明它们是否保持物体。我们引入了ManiPhysicsZoo,它将文献支持的材料属性、3D网格和支持性参考文献整合为可复用的物体资产。利用这些资产,一种基于求解器的评估方法根据物体几何形状、材料属性和记录的抓取条件计算特定于抓取的损伤阈值,并将其与记录的接触力进行比较,以评估潜在的变形和断裂。基于这些组件,ManiPhysicsBench在LIBERO和SimplerEnv中评估了三个物理轴和三个难度级别下的物体保持情况。公开的VLA检查点显示任务成功与安全成功(定义为完成任务同时保持物体)之间存在显著差距。它们的夹爪命令集中在完全打开和关闭附近,各物体之间的聚合分布大体相似,这与二值夹爪监督一致。我们通过重新训练VLA模型来检查物体特定的连续夹爪标签如何改变模型行为。重新训练的模型表现出更多依赖物体的抓取和更高的安全成功,但任务成功较低,且物体保持行为的泛化能力有限。

英文摘要

Vision-language-action (VLA) models aim to perform diverse manipulation tasks, but task success in existing rigid-body benchmarks does not indicate whether they preserve objects. We introduce ManiPhysicsZoo, which consolidates literature-supported material properties, 3D meshes, and supporting references into reusable object assets. Using these assets, a solver-based assessment computes grasp-specific damage thresholds from object geometry, material properties, and recorded grasp conditions and compares them with recorded contact forces to assess potential deformation and fracture. Building on these components, ManiPhysicsBench evaluates object preservation in LIBERO and SimplerEnv across three physics axes and three difficulty levels. Public VLA checkpoints show a substantial gap between task success and safe success, defined as task completion while preserving the object. Their gripper commands concentrate near full opening and closure, with largely similar aggregate distributions across objects, consistent with binary gripper supervision. We examine how object-specific continuous gripper labels change model behavior by retraining a VLA model. The retrained model shows more object-dependent gripping and higher safe success, but lower task success and limited generalization of object-preserving behavior.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑