SoftVTBench:面向可变形物体操作的形变感知视觉-触觉数据集与基准
SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation
浏览论文内容
中文总结 AI 辅助
本研究推出SoftVTBench视觉-触觉数据集与基准,定义形变感知成功率(DSR),发现触觉信息本身未必提升多模态融合,为可变形物体操作的物理交互研究提供资源。
中文摘要 AI 辅助
物理交互质量是可变形物体操作的核心,但大多数基准仅评估任务成功与否。策略可能在完成任务的同时出现滑动或过度压缩。主要瓶颈在于缺乏视觉-触觉数据集,该数据集需将策略可见的接触观测与完整任务上的独立物理真值配对。我们推出SoftVTBench,这是一个面向物理交互感知的可变形物体操作的视觉-触觉数据集,包含4000个专家演示和50多个资产,包括体积可变形物体和视觉匹配的刚性孪生体。每集以20Hz频率同步多视角RGB、双指触觉RGB、标记运动、本体感觉、语言、二元及连续夹爪动作,以及仅评估者可用的有限元(FEM)状态。基于该数据集,我们建立了一个闭环基准,使用固定的对象特定校准定义形变感知成功率(DSR),仅当执行完成任务且峰值归一化形变在容差范围内时,才将其计为成功。在扩散策略、π₀.₅和FastWAM中,所有12个分布内配置都包含违反形变容差的成功执行,占每个配置成功数的0.7%至24%。在分布偏移下,视觉-触觉变体在全部6个策略套件比较中实现更高的任务成功率,且在5个比较中实现更高的DSR,而它们在分布内的益处则参差不齐。这些结果表明,提供触觉信息本身并不能确保有效的多模态融合。因此,SoftVTBench提供了一个通用的视觉-触觉资源,不仅用于研究策略是否成功,还用于研究其与可变形物体的物理交互方式,以及触觉何时能改善这种交互。
英文摘要
Physical interaction quality is central to deformable-object manipulation, yet most benchmarks evaluate task success alone. A policy may complete the task while allowing slip or causing excessive compression. A primary bottleneck is the absence of visuo-tactile datasets that pair policy-visible contact observations with independent physical ground truth over complete tasks. We introduce SoftVTBench, a visuo-tactile dataset for physical-interaction-aware deformable-object manipulation. It contains 4,000 expert demonstrations and more than 50 assets, including volumetric deformable objects and visually matched rigid twins. At 20 Hz, each episode synchronizes multi-view RGB, dual-finger tactile RGB and marker motion, proprioception, language, and binary and continuous gripper actions, alongside evaluator-only finite-element (FEM) states. Building upon this dataset, we establish a closed-loop benchmark that uses fixed object-specific calibration to define the Deformation-aware Success Rate (DSR), which counts a rollout as successful only when it completes the task and keeps peak normalized deformation within tolerance. Across Diffusion Policy, $π_{0.5}$, and FastWAM, all 12 in-distribution configurations contain successful rollouts that violate the deformation tolerance, accounting for 0.7--24% of each configuration's successes. Under distribution shift, visuo-tactile variants achieve higher task success in all six policy--suite comparisons and higher DSR in five, whereas their in-distribution benefits are mixed. These results show that making touch available does not by itself ensure effective multimodal fusion. SoftVTBench therefore provides a common visuo-tactile resource for studying not only whether a policy succeeds, but how it physically interacts with deformable objects and when touch improves that interaction.
发表机构
- Tuojing Intelligence(拓境智能)
- Tsinghua University(清华大学)
- Southeast University(东南大学)
- Stevens Institute of Technology(斯蒂文斯理工学院)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- University of Manchester(曼彻斯特大学)
- Simple AI
- Imperial College London(帝国理工学院)
- Carnegie Mellon University(卡内基梅隆大学)
- Zhejiang University(浙江大学)
- Beihang University(北京航空航天大学)
- The University of Hong Kong(香港大学)
机构由 AI 辅助整理,请以论文原文为准。