arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GaussianWAM:将三维高斯场的几何与语义知识蒸馏至世界-动作模型

GaussianWAM: Distilling Geometry and Semantics from 3D Gaussian Fields into World-Action Models

Zijian Zhang, Yuqing Jiang, Weitao Zhou, Minglei Li, Jinhao Zhang, Yao Mu, Xiaofan Li, Hao Zhao, Haibao Yu

arXiv 2608.24714首次发表:更新:

发表机构

Tuojing Intelligence; University of Chinese Academy of Sciences; Institute of Automation, Chinese Academy of Sciences; Tsinghua University; Simple AI; Harbin Institute of Technology (Shenzhen); Shanghai Jiao Tong University; Zhejiang University; Institute for AI Industry Research (AIR), Tsinghua University; The University of Hong Kong(拓境智能; 中国科学院大学; 中国科学院自动化研究所; 清华大学; 简智人工智能; 哈尔滨工业大学(深圳); 上海交通大学; 浙江大学; 清华大学人工智能产业研究院; 香港大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

GaussianWAM是训练时的表征增强框架,通过三维高斯场蒸馏几何与语义监督,在LIBERO-Plus等任务上显著提升了WAM类模型的机器人操作性能,且不改变部署架构。

AI 中文摘要

世界-动作模型(World-Action Models, WAMs)联合学习未来视觉预测与动作生成,利用视频动力学作为机器人操作的表征学习信号。然而,它们的视频隐层主要针对视觉预测进行优化,未明确鼓励保留跨视图几何结构或空间定位的、与物体相关的语义。我们提出GaussianWAM,这是一种训练时的表征增强框架,通过三维高斯场组织几何与语义监督信号。给定同步的多视图观测,冻结的几何与视觉基础模型提供深度、相机参数及密集语义特征。GaussianWAM将这些异构信号绑定至共享高斯基元,并渲染空间对齐的语义、深度与覆盖度目标,这些目标被蒸馏至WAM的当前观测表征中。训练后,所有教师模型、高斯组件及辅助预测头均被移除,仅保留原始WAM推理路径,无额外模块或前向计算。在LIBERO-Plus上,GaussianWAM将FastWAM的性能从52.05%提升至71.29%,将Cosmos Policy从71.52%提升至77.30%。直接CLIP和VGGT蒸馏已为FastWAM建立69.37%的强基线,而高斯场统一进一步将其提升至71.29%,证明了空间组织异构教师信号的益处。GaussianWAM还提升了标准LIBERO上的性能,并在RoboTwin及真实世界操作任务上呈现正向迁移趋势。这些结果表明,训练时的高斯蒸馏为向WAM表征注入几何与语义相关监督提供了实用方法,且无需改变其部署架构。

英文摘要

World-Action Models (WAMs) jointly learn future visual prediction and action generation, using video dynamics as a representation-learning signal for robotic manipulation. However, their video latents are primarily optimized for visual prediction and are not explicitly encouraged to preserve cross-view geometric structure or spatially localized, object-relevant semantics. We propose \textbf{GaussianWAM}, a training-time representation-enhancement framework that organizes geometric and semantic supervision through a 3D Gaussian field. Given synchronized multi-view observations, frozen geometry and vision foundation models provide depth, camera parameters, and dense semantic features. GaussianWAM binds these heterogeneous signals to shared Gaussian primitives and renders spatially aligned semantic, depth, and coverage targets, which are distilled into the current-observation representations of the WAM. All teacher models, Gaussian components, and auxiliary prediction heads are removed after training, leaving the original WAM inference path without additional modules or forward computation. On LIBERO-Plus, GaussianWAM improves FastWAM from 52.05\% to 71.29\% and Cosmos Policy from 71.52\% to 77.30\%. Direct CLIP and VGGT distillation already establishes a strong FastWAM baseline of 69.37\%, while Gaussian-field unification further improves it to 71.29\%, supporting the benefit of spatially organizing heterogeneous teacher signals. GaussianWAM also improves performance on standard LIBERO and shows positive transfer trends on RoboTwin and real-world manipulation. These results suggest that training-time Gaussian distillation provides a practical way to inject geometry- and semantics-related supervision into WAM representations without changing their deployment architecture.

Comments13 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑