arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过本地编码智能体泛化操作技能

Generalizing Manipulation Skills with a Local Coding Agent

Raman Talwar, Elias Nijs, Andreas Verleysen, Francis wyffels

arXiv 2609.26499首次发表:更新:

发表机构

Ghent University – imec(根特大学 – imec)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究本地开放权重视觉-语言模型通过编码智能体控制机械臂,无需新编程或训练即可一次性泛化到新任务,在45次试验中30次成功,并发现重做任务时间减少50%,表明存在自我改进能力。

AI 中文摘要

如今,开放权重语言模型的进展使得系统能够在单台工作站上编写、执行和调试代码。大多数语言驱动的机器人给模型一个固定的动作接口或训练好的策略。因此,泛化到新任务意味着更多的工程努力或更多的数据收集,两者都耗时。我们研究本地开放权重视觉-语言模型能否控制机器人,并在无需新的人类编程或训练的情况下,一次性泛化到任务的新变体。我们让本地开放权重VLM,Qwen3.8-27B,通过编码智能体框架驱动UR3e机械臂。它在一个实现运动学、安全限制和经典计算机视觉技术的服务之上编写并运行自己的代码。我们研究该系统能否泛化到未见过的任务。具体来说,我们在九个由儿童玩具构建的任务上测试它,这些任务旨在探测跨各种物体特征的泛化能力:颜色、大小、形状以及这些特征的任务变体。每个任务进行五次试验,我们在45次试验中观察到30次泛化成功,持续时间根据任务复杂度从3.4分钟到67.5分钟不等。我们进一步测试在成功完成后要求智能体重做任务是否有加速。这导致持续时间减少50%,表明存在随时间的自我改进。最后,我们揭示了本地编码智能体的局限性。我们相信解决这些局限性并结合对自我改进的进一步研究,指向了本地编码智能体实际部署的直接路径。

英文摘要

Today, progress in open-weight language models enables systems capable of writing, executing and debugging code while still running on a single workstation. Most language-driven robots give the model a fixed action interface or a trained policy. Generalizing to a new task therefore means more engineering effort or more data collection, both time-consuming. We investigate whether a local open-weight vision-language model can control a robot and one-shot generalize to new variations of a task without new human programming or training. We let a local open-weight VLM, Qwen3.8-27B, drive a UR3e robotic arm from a coding-agent harness. It writes and runs its own code above a service that implements kinematics, safety limits and classic computer vision techniques. We investigate if this system is capable of generalizing to unseen tasks. Specifically, we test it on nine tasks built from children's toys designed to probe generalization capability across various object characteristics: color, size, shape, and task variation of those. With five trials for each task, we observe generalization in 30 out of 45 trials with durations ranging from 3.4 to 67.5 minutes depending on task complexity. We further test if there is a speedup when an agent is asked to redo the task after successful completion. This resulted in a 50% reduction in duration, indicating that there is self-improvement over time. Finally, we expose the limitations of a local coding agent. We believe that solving those limitations combined with further investigation of self-improvement over time points at a direct path toward real-world deployment of a local coding agent.

Comments8 pages, 4 figures, 5 tables. Raman Talwar and Elias Nijs contributed equally

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑