arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于高斯泼溅的单次扫描演示合成用于视觉运动策略学习

Demonstration Synthesis from a Single Scan via Gaussian Splatting for Visuomotor Policy Learning

Beichen Wang, Yuen-Hei Yeung, V. R. Sridhar Devarakonda, Xuesu Xiao

arXiv 2609.21112首次发表:更新:

发表机构

George Mason University; New York University(乔治梅森大学; 纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出GaussianFactory,通过单次视频扫描重建3D高斯场景并运动学规划抓取轨迹,无需物理引擎即可合成高保真演示,训练扩散策略在仿真和真实场景分别达到95.1%和84.2%成功率。

AI 中文摘要

训练视觉运动策略需要大量与目标环境高度匹配的演示,但重新收集这些演示的成本仍然很高。现有的演示合成方法降低了这一成本,但仍受限于高人工投入、有限的视觉保真度或对物理模拟器的严重依赖。本文介绍了GaussianFactory,一个高保真数据引擎,仅需一次视频扫描作为唯一人工输入,且生成过程中无需物理引擎,即可批量生产演示。具体而言,GaussianFactory将场景重建为可编辑的3D高斯泼溅(3DGS)副本,并从场景支持的对象组合任务中采样。对于每个任务,它纯粹在几何重建上进行运动学层面的抓取和轨迹规划,渲染出与目标环境视觉匹配的照片级逼真演示。物理动力学仅在接触力交互决定结果时进入流程——即在抓取形成阶段,通过一个在交互数据集上预训练一次的学习接触模型。为评估合成演示的下游实用性,我们在两种设置中实现了端到端的扫描到部署工作流:一个模拟真实世界的仿真场景以确保可复现性,以及一个配备物理UR10e机器人的真实工作空间。在每种设置中,仅使用合成演示训练的标准扩散策略分别达到了95.1%和84.2%的成功率。

英文摘要

Training a visuomotor policy calls for abundant demonstrations that closely match the target environment, yet collecting them anew remains expensive. Existing demonstration synthesis methods reduce this cost but remain constrained by high manual effort, limited visual fidelity, or heavy reliance on physics simulators. This paper introduces GaussianFactory, a high-fidelity data engine that mass-produces demonstrations with a single video scan as its only human input and no physics engine in the generation loop. Specifically, GaussianFactory reconstructs the scene as an editable 3D Gaussian Splatting (3DGS) replica and samples from the object-combination tasks the scene affords. For each task, it plans grasps and trajectories purely kinematically on the geometric reconstruction, rendering photorealistic demonstrations that visually match the target environment. Physical dynamics enter the pipeline only where contact force interactions dictate the outcome---during grasp formation, via a learned contact model pretrained once on an interaction dataset. To evaluate the downstream utility of the synthesized demonstrations, we implement an end-to-end scan-to-deployment workflow in two setups: a simulated scene that stands in for the real world to enable reproducibility, and a real-world workspace with a physical UR10e robot. In each setup, a standard diffusion policy trained solely on the synthesized demonstrations achieves 95.1% and 84.2% success rates, respectively.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑