arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.01856cs.RO

ChunkVLA-AM:增材制造中视觉-语言-动作机器人控制的并行动作分块

ChunkVLA-AM: Parallel Action Chunking for Vision-Language-Action Robot Control in Additive Manufacturing

Zhugang Liu, Kaichuang Zhang, Jinman Zhang, Pu Sun, Martha Asare, Jose Hernandez, Maxim Ermolinsky, Efren Saenz, Qi Lu, Jinghao Yang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出在增材制造工作单元中将OpenVLA-OFT部署到FAIRINO FR3机器人的框架,通过数据管道和并行动作分块实现闭环控制,在42次试验中达到92.9%成功率。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型统一了视觉感知、语言理解和动作生成,为增材制造(AM)中的自动化提供了新的机遇。然而,在AM中的部署仍具挑战性,因为将这些模型适配到未见过的机器人本体上成本高昂,且性能在环境变化时可能下降。在这项工作中,我们提出了一个框架,用于在固定的AM工作单元中将OpenVLA-OFT部署到FAIRINO FR3机器人上。一个数据管道将单目真实世界演示转换为与OpenVLA兼容的TFDS/RLDS数据集,以支持对FR3本体的适配。在运行时,每个推理请求预测一个包含8步的7维动作块。FR3以开环方式执行每个动作块,然后捕获新的观测,从而在块之间提供闭环反馈。该系统采用云-边架构,其中FR3客户端通过FastAPI接口将观测流传输到远程推理服务器。在42次物理A到B物体转移试验中,红蓝目标各占一半,系统成功39次(92.9%)。所有三次失败均发生在最终放置阶段,原因是释放高度控制不足导致物体倾倒。一次光照扫描确定了在0-255标度上低误差的亮度范围为85-125,其中最低平均空间误差出现在95。

英文摘要

Vision-language-action (VLA) models unify visual perception, language understanding, and action generation, offering new opportunities for automation in additive manufacturing (AM). However, deployment in AM remains challenging because adapting these models to unseen robot embodiments is costly, and performance can degrade under environment changes. In this work, we present a framework for deploying OpenVLA-OFT on a FAIRINO FR3 robot in a fixed AM workcell. A data pipeline converts monocular real-world demonstrations into OpenVLA-compatible TFDS/RLDS datasets to support adaptation to the FR3 embodiment. At runtime, each inference request predicts an eight-step chunk of 7-D actions. The FR3 executes each chunk open loop before capturing a new observation, providing closed-loop feedback between chunks. The system uses a cloud-edge architecture in which the FR3 client streams observations to a remote inference server through a FastAPI interface. In 42 physical A-to-B object-transfer trials, evenly split between red and blue targets, the system succeeded in 39 (92.9%). All three failures occurred during final placement, when insufficient release-height control caused the object to topple. An illumination sweep identified a low-error luminance range of 85-125 on a 0-255 scale, with the lowest mean spatial error at 95.

发表机构

  • The University of Texas Rio Grande Valley(德克萨斯大学里奥格兰德河谷分校)
  • University of South Florida(南佛罗里达大学)
  • San Diego State University(圣地亚哥州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑