HEARTH:面向温度感知机器人操作的对象中心RGB-热-3D数据集
HEARTH: An Object-Centric RGB-Thermal-3D Dataset for Temperature-Aware Robot Manipulation
- Simon Fraser University(西蒙菲莎大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出HEARTH,一个包含90个日常对象、145种状态的对象中心RGB-热-3D数据集,通过将温度映射到网格并微调VLA模型,使温度相关对象选择任务成功率从35.0%提升至75.0%。
AI中文摘要:
语言引导的操作可能依赖于可见外观无法揭示的物理属性。温度就是其中之一,但用于机器人学习的对象数据集很少将测量温度与对象外观和几何形状关联起来。我们提出了HEARTH,一个对象中心的RGB-热-3D数据集,包含来自18个日常类别的90个物理对象,共145个捕获的对象状态。我们的流程通过相机标定和姿态传递将表观表面温度映射到重建的网格上。该数据集包括原始温度测量值、相机参数、RGB纹理网格和用于仿真的热纹理。我们利用这些资源构建了三个源自LIBERO的任务,并收集了1,200个演示,用于微调预训练的视觉-语言-动作(VLA)模型$\pi_{0.5}$。在消融研究中,向VLA添加热观测将温度相关对象选择任务的成功率从仅RGB基线的35.0%提高到75.0%。这些结果证明了HEARTH在训练机器人策略遵循温度相关指令方面的实用性。
英文摘要:
Language-guided manipulation can depend on physical properties that visible appearance does not reveal. Temperature is one such property, but object datasets for robot learning rarely associate measured temperatures with object appearance and geometry. We present HEARTH, an object-centric RGB-thermal-3D dataset of 90 physical objects from 18 everyday categories, comprising 145 captured object states. Our pipeline maps apparent surface temperatures onto reconstructed meshes through camera calibration and pose transfer. The dataset includes raw temperature measurements, camera parameters, RGB-textured meshes, and thermal textures for simulation. We use these assets to construct three LIBERO-derived tasks and collect 1,200 demonstrations for fine-tuning a pretrained vision-language-action (VLA) model, $π_{0.5}$. In an ablation study, adding thermal observations to the VLA increases success on temperature-dependent object-selection tasks from 35.0% for the RGB-only baseline to 75.0%. These results demonstrate the utility of HEARTH for training robot policies to follow temperature-related instructions.