发表机构
University of Augsburg(奥格斯堡大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将SmolVLA适配于Universal Robots轻量机器人,发布ROS2SmolVLA开源接口,通过UR10e取放任务验证其功能,为工业级轻量机器人提供了本地可计算的小型VLA模型方案。
AI 中文摘要
工业需求的变化改变了生产范式,由于批量更小、产品种类更多,企业在采用更具适应性的生产系统时面临日益严峻的挑战。特别是基于机器人的自动化通常是静态的,无法响应不断变化的流程。视觉-语言-动作(Vision-Language-Action,VLA)模型为缓解这一挑战提供了有前景的机会,它可基于观测到的系统状态生成机器人动作。然而,当前研究要么聚焦于无法在本地计算的大型模型,带来合规性和安全性挑战,要么使用实验室级机器人硬件,阻碍了其在实际工业场景中的应用。本研究将Hugging Face的SmolVLA适配于Universal Robots轻量机器人,还发布了开源仓库ROS2SmolVLA,该仓库实现了ROS 2与SmolVLA的接口,使其可应用于工业级硬件,从而便于在实验室和工业环境中灵活采用。我们通过取放任务验证了SmolVLA在Universal Robots UR10e上的功能,并给出了实施指南。研究结果表明,SmolVLA是适合需要本地计算的小型任务的合适选项。
英文摘要
Industrial demand changes the paradigms of production. Due to smaller batch sizes and more variations in products, companies face a growing challenge to adopt more adaptive production systems. In particular, robot-based automation is usually static and fails to respond to constantly changing processes. Vision-Language-Action (VLA) Models are a promising opportunity to mitigate this challenge by generating robot actions based on the observed system state. However, current research either focuses on large models that cannot be computed on premise, creating compliance and security challenges, or use lab-grade robot hardware that obscures exploitation in real industrial settings. In this work, we adapt Hugging Face's SmolVLA for Universal Robots lightweight robots. Further, we release the open-source repository ROS2SmolVLA that implements an interface for ROS 2 to SmolVLA, and makes it applicable for industrial-grade hardware. By this, we allow a lenient adoption into lab and industrial environments. We validate the functionality of SmolVLA for a Universal Robots UR10e using a pick-and-place task and give implementation guidelines. Our findings support that SmolVLA is a well-suited option for small-sized tasks that need to be computed on premise.
CommentsAccepted at 8th International Conference on Industry of the Future and Smart Manufacturing, Padua & Venice, 2026