发表机构
KTH; EPFL; Idiap Research Institute; Beihang University; The Hong Kong University of Science and Technology (Guangzhou); Great Bay University; The Hong Kong Polytechnic University(瑞典皇家理工学院; 洛桑联邦理工学院; Idiap 研究所; 北京航空航天大学; 香港科技大学(广州); 大湾区大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对机器人长程操作中执行错误累积问题,提出Robot-GST框架,利用3D高斯泼溅和SAM3D构建高保真环境,结合时空推理与状态感知执行,实现先模拟评估后行动,提升操作可靠性。
AI 中文摘要
随着机器人操作策略越来越依赖视觉语言模型进行端到端决策,其发展迅速。然而,可靠部署仍然具有挑战性,因为许多策略缺乏预测任务结果和评估生成动作是否能实现期望最终状态的显式机制,导致在长程操作中执行错误累积。我们提出了Robot-GST,一个几何感知的时空行为表示与评估框架,该框架构建了高斯-SAM机器人环境用于真实到模拟的策略验证,并提高了真实世界操作部署的可靠性。我们的方法使用3D高斯泼溅和SAM3D从RGB-D观测构建高保真机器人环境,实现“先模拟评估后行动”。它集成视觉观测和语言指令,利用大型视觉语言模型进行长程任务规划的时空推理。为了弥合高层规划与真实世界执行之间的鸿沟,我们通过几何采样和基于状态的轨迹规划引入了高斯感知的最终状态估计。在执行前,候选动作序列在高斯-SAM环境中进行模拟和评估,以过滤不可行的行为。我们在涉及刚体、软体和可变形物体的代表性操作任务上验证了我们的方法,包括立方体放置、玩具打包和鸭子重排,证明了几何感知的时空推理和状态感知执行提高了不同物体类别下的操作可靠性。我们的结果表明,将几何感知重建与高质量渲染和模拟相结合,为评估机器人操作行为提供了一种可扩展的方法。网站:此HTTPS URL。
英文摘要
Robotic manipulation policies are advancing rapidly with increasing reliance on vision-language models for end-to-end decision making. However, reliable deployment remains challenging because many policies lack explicit mechanisms for predicting task outcomes and evaluating whether generated actions will achieve desired final states, causing execution errors to accumulate during long-horizon manipulation. We present Robot-GST, a geometry-aware spatio-temporal behaviour representation and evaluation framework that constructs a Gaussian-SAM robotic environment for real-to-sim policy verification and improves the reliability of real-world manipulation deployment. Our approach constructs a high-fidelity robotic environment from RGB-D observations using 3D Gaussian Splatting and SAM3D, enabling ``simulation and evaluation before acting''. It integrates visual observations and language instructions with spatio-temporal reasoning for long-horizon task planning using large vision-language models. To bridge high-level planning and real-world execution, we introduce Gaussian-aware final-state estimation through geometric sampling and state-based trajectory planning. Before execution, candidate action sequences are simulated and evaluated in the Gaussian-SAM environment to filter infeasible behaviours. We validate our approach on representative manipulation tasks involving rigid, soft, and deformable objects, including cube placing, toy packing, and duck rearrangement, demonstrating that geometry-aware spatio-temporal reasoning and state-aware execution improve manipulation reliability across different object categories. Our results suggest that combining geometry-aware reconstruction with high-quality rendering and simulation provides a scalable approach for evaluating robotic manipulation behaviours. Website: https://robot-gst.github.io
Comments9 pages, 8 figures, 2 tables