arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

世界到手腕:面向精细机器人操作的任务条件未来手腕建模

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

Yuhao Pan, Haosong Peng, Zhengshen Zhang, Zhengyang Yan, Yalun Dai, Fushuo Huo, Chujie Wang, Tianyu Qi, Xiucheng Wang, Nan Cheng, Wenchao Xu

arXiv 2608.05369首次发表:更新:

发表机构

The Hong Kong University of Science and Technology; National University of Singapore; Nanyang Technological University; Wuhan University; Sun Yat-sen University; Xidian University; Southeast University(香港科技大学; 新加坡国立大学; 南洋理工大学; 武汉大学; 中山大学; 西安电子科技大学; 东南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出W2-VLA模型,通过任务条件未来手腕建模结合W2-CoT辅助监督,在LIBERO等数据集及真实任务中提升了机器人精细接触敏感操作能力,动作生成速率超80Hz。

AI 中文摘要

视觉-语言-动作(VLA)模型通常将主视角和手腕视角观测视为并行视觉输入,忽略了它们在机器人操作中的不同作用。然而,精细操作需要预判在全局任务语境下,手腕局部交互将如何演变。为解决这一局限,我们提出World-to-Wrist VLA(W2-VLA),这是一种用于精细机器人操作的VLA模型,具备任务条件未来手腕建模能力。给定当前多视角观测和任务指令,W2-VLA将一组潜在建模词元上下文化为视觉-语言模型与手腕预测器之间的紧凑接口。基于该接口和观测到的手腕历史,预测器会预测未来手腕潜在变量,这些变量被转换为具有未来感知的上下文,用于动作预测。此外,我们引入W2-CoT,这是一种生成结构化标注的合成流水线,该标注描述操作进度、物理过渡线索和手腕局部证据,这些标注提供辅助监督以塑造任务条件潜在接口。在LIBERO、RoboTwin 2.0和真实世界操作任务上的实验表明,该模型在单臂和双臂设置下均提升了精细且接触敏感的操作能力,同时保持动作生成速率高于80 Hz。

英文摘要

Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, LIBERO-Plus, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across single-arm and bimanual settings, while maintaining real-time action generation above $80$~Hz.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑