arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31717cs.CVcs.ROeess.IV

PanoFuse:基于解耦语义-几何路由的全景增强视觉-语言-动作学习

PanoFuse: Panorama-Enhanced Vision-Language-Action Learning with Decoupled Semantic-Geometric Routing

Peng Xu, Haoran Lin, Wanjun Jia, Kai Luo, Wenrui Chen, Zhiyong Li, Kailun Yang

首次发表
浏览论文内容

中文总结 AI 辅助

针对VLA策略因透视相机视场有限而缺乏全局上下文的问题,提出PanoFuse全景增强框架,通过解耦语义-几何路由融合全景感知,在七项评估中平均成功率52.9%,超越基线。

中文摘要 AI 辅助

视觉-语言-动作(VLA)策略在语言条件下的机器人操作任务中展现出良好的性能。然而,现有的大多数VLA系统依赖于视场有限的传统透视相机,往往缺乏全局场景上下文,在视觉遮挡、干扰物和未见环境中会导致操作不可靠。在这项工作中,我们提出了PanoFuse,一种全景增强的VLA框架,用全局全景感知补充局部操作观察。PanoFuse引入了一个专门的全景分支,利用预训练的全景基础模型从全向观察中提取互补的语义和几何表示。我们不是直接混合这些异构特征,而是引入了解耦语义-几何路由(DSGR),将语义和几何表示作为独立的上下文流,并通过结构化的分块注意力将两者选择性地路由到下游的状态和动作表示中。这种设计为动作专家提供了全局空间上下文,同时保留了来自预训练VLA骨干的任务相关语义信息。我们进一步开发了同步数据采集流程,并构建了一个新的真实世界操作数据集,包含全景RGB观测、腕部视图图像、语言指令、机器人状态和动作。在七个评估设置中,PanoFuse的平均成功率达到52.9%,优于评估的基线方法,并在新物体、未见背景和干扰物丰富的设置中持续获得性能提升。代码和数据将在该https URL上公开发布。

英文摘要

Vision-Language-Action (VLA) policies have shown promising performance in language-conditioned robotic manipulation. However, most existing VLA systems rely on conventional perspective cameras with limited fields of view, often missing global scene context and leading to unreliable manipulation under visual occlusions, distractors, and unseen environments. In this work, we propose PanoFuse, a panorama-enhanced VLA framework that complements local manipulation observations with global panoramic perception. PanoFuse introduces a dedicated panoramic branch that leverages a pretrained panoramic foundation model to extract complementary semantic and geometric representations from omnidirectional observations. Rather than directly mixing these heterogeneous features, we introduce Decoupled Semantic-Geometric Routing (DSGR), which maintains semantic and geometric representations as separate context streams and selectively routes both to downstream state and action representations through structured block-wise attention. This design provides the action expert with global spatial context while preserving task-relevant semantic information from the pretrained VLA backbone. We further develop a synchronized data collection pipeline and construct a new real-world manipulation dataset containing panoramic RGB observations, wrist-view images, language instructions, robot states, and actions. Across seven evaluation settings, PanoFuse achieves an average success rate of 52.9%, outperforming the evaluated baselines and achieving consistent gains under novel-object, unseen-background, and distractor-rich settings. Code and data will be released publicly at https://xux-hnu.github.io/PanoFuse.

补充信息

↑