arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29389cs.RO

Robo-Harness K1:通过感知增强利用机器人使用智能体

Robo-Harness K1: Harnessing Robot-Use Agents via Perception Augmentation

Zexi Li, Yehang Zhang, Wenqian Li, Haojian Huang, Chenxu Wang, Shiyuan Deng, Yangkai Wei, Tianyi Zhang, Binghui Xie, Bohan Zhou, Yifan Chang, Kaiwen Zhou, Ying-… 展开作者

Zexi Li, Yehang Zhang, Wenqian Li, Haojian Huang, Chenxu Wang, Shiyuan Deng, Yangkai Wei, Tianyi Zhang, Binghui Xie, Bohan Zhou, Yifan Chang, Kaiwen Zhou, Ying-Cong Chen, James Cheng, Yinchuan Li

AI总结:

提出Robo-Harness K1框架,将感知作为工具暴露给VLM智能体,无需修改架构或训练深度编码器,在LIBERO-PRO和RoboTwin任务上显著提升准确率,并展示了样本高效和可泛化性。

AI中文摘要:

基础视觉语言模型(VLM)能够理解物体、指令和空间关系,但将这种能力转化为机器人操作仍然困难。视觉-语言-动作(VLA)模型需要大量演示,并可能损害预训练的理解能力,而直接基于RGB的VLM控制成本高昂且严重依赖模型能力。我们提出了Robo-Harness K1,一个将感知暴露为工具的机器人使用智能体(RUA)框架。该智能体查询校准深度、持久视觉锚点、空间测量和抓取假设,然后根据返回的证据选择通用动作。这种接口在不改变VLM架构或训练深度编码器的情况下使3D几何信息可访问。在匹配的LIBERO-PRO任务上,配备K1的Gemini 3.7 Flash达到77.8%的准确率,超过了使用仅RGB工具的GPT-6 Astra的61.1%;K1进一步将Astra提升至88.9%。无需目标微调,配备K1的Gemini可迁移到三个RoboSuite机械臂和双臂RoboTwin任务。在RoboTwin上,它在Easy任务上达到32.0%,在Hard任务上达到28.0%,显示出对视觉和环境扰动的鲁棒性。K1还产生与下一标记训练对齐的工具调用轨迹。一个仅在107个教师片段上训练的Qwen3.5-9B学生模型在新初始状态上达到44.2%的准确率,而OpenVLA为30.2%;在保留任务条件下达到13.9%,而OpenVLA为0.0%。这些结果表明,感知增强的RUA为样本高效、可泛化的机器人策略提供了一条有前景的途径,通过可访问的工具接口利用VLM能力。

英文摘要:

Foundation vision-language models (VLMs) understand objects, instructions, and spatial relations, yet translating this capability into robotic manipulation remains difficult. Vision-language-action (VLA) models require extensive demonstrations and may compromise pretrained understanding, while direct RGB-only VLM control is costly and strongly dependent on model capability. We introduce Robo-Harness K1, a robot-use agent (RUA) framework that exposes perception as tools. The agent queries calibrated depth, persistent visual anchors, spatial measurements, and grasp hypotheses, then selects generic motions from the returned evidence. This interface makes 3D geometry accessible without changing the VLM architecture or training a depth encoder. On matched LIBERO-PRO tasks, Gemini 3.7 Flash with K1 reaches 77.8% accuracy, surpassing GPT-6 Astra's 61.1% with an RGB-only harness; K1 further improves Astra to 88.9%. Without target fine-tuning, Gemini with K1 transfers to three RoboSuite arms and dual-arm RoboTwin tasks. On RoboTwin, it achieves 32.0% on Easy and 28.0% on Hard, showing resilience to visual and environmental perturbations. K1 also produces tool-call traces aligned with next-token training. A Qwen3.5-9B student trained on only 107 teacher episodes reaches 44.2% accuracy on new initial states versus 30.2% for OpenVLA, and 13.9% on held-out task conditions versus 0.0% for OpenVLA. These results suggest that perception-augmented RUAs offer a promising route to sample-efficient, generalizable robotic policies that leverage VLM capabilities through an accessible tool interface.

补充信息

↑