arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17453cs.RO

EATR-Stereo:面向人形机器人视觉-语言-动作控制的具身感知配对立体证据路由

EATR-Stereo: Embodiment-Aware Token Routing of Paired Stereo Evidence for Humanoid Vision-Language-Action Control

  • Harbin Institute of Technology(哈尔滨工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Songwei Wu, Rui Zhao, Fan Yang, Zhongqiang Nie, Zhiduo Jiang, Wandong Sun, Yuwei Li, Jian Hu, Yang Liu, Hong Liu

中文总结 AI 辅助

EATR-Stereo是具身感知的令牌路由框架,在冻结预训练VLA模型的前提下,通过选择性路由配对立体证据,提升了人形机器人长时程视觉-语言-动作控制的任务成功率。

中文摘要 AI 辅助

头戴式立体相机的长时程人形机器人视觉-语言-动作(VLA)控制,需要能利用互补视角同时兼容预训练表征的视觉接口。现有接口常丢弃互补立体证据,或融合额外观测却未保留原生主视角通路,也未让辅助信息适配机器人具身。我们提出EATR-Stereo,一种具身感知的令牌路由框架,该框架保留主视角令牌,通过查询同步的辅助视角令牌序列构建主视角对齐的跨视角辅助令牌(CVATs)。分段式本体感受编码器进一步基于机器人构型历史,对令牌级辅助使用进行条件约束,使动作生成过程中能选择性融入立体证据。路由后的辅助流在冻结预训练VLA的视觉-语言模型的同时,增强其语言和主视觉上下文。我们在具有37维本体感受状态的33自由度物理人形机器人上,对9种配置开展了时长超100秒的搜索-接近-抓取-放置-返回任务评估。EATR-Stereo实现了60.0%的全任务成功率、100.0%的抓取成功率、80.0%的阶段成功率;在严重非对称遮挡下,其恢复性能提升至80%,而仅使用CVAT的方法仅为30%。消融研究进一步表明,保留主视角令牌、结合跨视角辅助特征与结构化本体感受路由具有重要意义。这些结果表明,选择性路由的配对立体证据可提升空间 grounding,助力可靠的长时程人形机器人VLA控制。

英文摘要

Long-horizon humanoid vision--language--action (VLA) control with head-mounted stereo cameras requires visual interfaces that can exploit complementary views while maintaining compatibility with pretrained representations. Existing interfaces often discard complementary stereo evidence or fuse additional observations without preserving the native primary-view pathway and adapting auxiliary information to robot embodiment. We present EATR-Stereo, an embodiment-aware token-routing framework that retains primary-view tokens and constructs primary-aligned Cross-View Auxiliary Tokens (CVATs) by querying the synchronized auxiliary-view token sequence. A body-segmented proprioceptive encoder further conditions token-wise auxiliary usage on robot configuration history, enabling selective incorporation of stereo evidence during action generation. The routed auxiliary stream augments the language and primary-visual context of a pretrained VLA while keeping its vision--language model frozen. On a 33-DoF physical humanoid with a 37-D proprioceptive state, we evaluate nine configurations in over-100-s search--approach--grasp--place--return tasks. EATR-Stereo achieves 60.0% full-task success, 100.0% grasp success, and 80.0% stage success. Under severe asymmetric occlusion, it improves recovery to 80% compared with 30% for CVAT alone. Ablation studies further show the importance of preserving primary tokens and combining cross-view auxiliary features with structured proprioceptive routing. These results demonstrate that selectively routed paired stereo evidence improves spatial grounding for reliable long-horizon humanoid VLA control.

补充信息

↑