arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15012cs.RO

原子运动坐标:用于语言可引导与力响应操作

Atomic Motion Coordinate for Language-Steerable and Force-Responsive Manipulation

Jiaqi Zhai, Jingkai Zhao, Chen Yang, Siyuan Ma, Yutian Zhang, Liwen Yang, Qinglian Wu, Weiqi Fan, Yifei Wang, Yi Zheng, Chenxi Gu, Dong Wei, Wei Zhang

AI总结:

提出原子运动坐标,通过几何坐标和力响应调制,实现语言引导的机器人操作,显著提升干预分离度和真实任务成功率。

AI中文摘要:

仅改变语言指令能否重定向VLA策略的末端执行器,还是视觉驱动的运动先验占主导?我们提出原子运动坐标(Atomic Motion Coordinate),一种基于几何的坐标,用于可引导和力响应的操作。每个机械臂拥有十三个带符号的平移、旋转和保持原子,这些原子从文本和前向运动学中提取,且不依赖视觉;该坐标通过加权码本对齐注入到每个动作专家模块中。接触历史通过一个有界的球面残差调制同一坐标,该残差从固定名义潜变量重新计算,仅再生未执行的水平线后缀。在7,520次离线水平线干预中,相反原子分离度达到92.5/83.1%(单/双臂),而LA4VLA风格仅为39.1/24.0%。在每个任务50次真实机器人试验中,AMC将OOD水果进度从60.5%提升至87.8%;力适应将插头/花瓶任务从59.0/71.5%提升至78.5/75.2%。

英文摘要:

Can changing only the language instruction redirect a VLA policy's end effector, or does the visually driven motion prior dominate? We present Atomic Motion Coordinate, a geometry-grounded coordinate for steerable and force-responsive manipulation. Each arm owns thirteen signed translation, rotation, and hold atoms grounded from text and forward kinematics with vision withheld, and the coordinate is injected into every action-expert block via weighted codebook alignment. Contact history modulates the same coordinate through a bounded spherical residual that is recomputed from a fixed nominal latent to regenerate only the unexecuted horizon suffix. Across 7,520 offline horizon interventions, opposite-atom separation reaches 92.5/83.1% (single/dual) versus 39.1/24.0% for LA4VLA-style. Across 50 real-robot trials per task, AMC raises OOD fruit progress from 60.5% to 87.8%; force adaptation raises Plug/Vase from 59.0/71.5% to 78.5/75.2%.

补充信息

↑