arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17210cs.ROcs.AI

FluxVLA引擎:具身智能的一站式VLA工程平台

FluxVLA Engine: A One-Stop VLA Engineering Platform for Embodied Intelligence

Yinhao Li, Weixin Mao, Zihan Lan, Jikun Rong, Qirui Hu, Yiming Zhang, Weipeng Deng, Bowen Shen, Minzhao Zhu, Yiming Mao, Yan Yang, Chenguang Cui, Hongyuan Chen,… 展开作者

Yinhao Li, Weixin Mao, Zihan Lan, Jikun Rong, Qirui Hu, Yiming Zhang, Weipeng Deng, Bowen Shen, Minzhao Zhu, Yiming Mao, Yan Yang, Chenguang Cui, Hongyuan Chen, Xu Huang, Zheyi Zhao, Pinxi Shen, Bozhen He, Zhen Fu, Yifan Wang, Zexin Zhang, Ang Gao, Haoyu Chen, Chengqi Shi, Hua Chen

首次发表
浏览论文内容

中文总结 AI 辅助

FluxVLA引擎是一个配置驱动的开源平台,通过标准化接口和集成工具链,将异构具身策略组件整合为可复现的数据到部署工作流,解决VLA模型工程化瓶颈。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型、世界-动作模型(WAMs)以及离线强化学习方法正在迅速扩展具身策略的设计空间,然而将这些算法转化为可靠的机器人系统仍受制于碎片化的数据格式、训练栈、评估协议、推理运行时以及具身特定接口。我们提出FluxVLA引擎,一个开放的、配置驱动的平台,它将异构的具身策略组件转化为可复现的数据到部署的工作流。该引擎并非引入另一种策略模型,而是标准化了数据集、视觉-语言与世界模型、动作头、奖励或优势加权学习、分布式训练、仿真评估、优化推理以及机器人操作器的接口。该引擎进一步集成了组合式双臂仿真、可扩展的自动数据生成,以及模型解耦的人在环回放、接管、修正收集和奖励标注。为了实现响应式的物理执行,它结合了实时分块(RTC)与加速推理后端、轻量级远程GPU服务以及可配置的轨迹后处理。这些能力共同通过共享且可审计的契约将离线学习、仿真验证、在线修正和真实机器人执行连接起来。因此,FluxVLA瞄准了将前景广阔的具身学习算法与可复现评估和可靠部署隔离开来的工程瓶颈。代码可在以下网址获取:https://this https URL

英文摘要

Vision-language-action (VLA) models, world-action models (WAMs), and offline reinforcement learning methods are rapidly expanding the design space of embodied policies, yet turning these algorithms into reliable robot systems remains constrained by fragmented data formats, training stacks, evaluation protocols, inference runtimes, and embodiment-specific interfaces. We present $\mathrm{FluxVLA}$ Engine, an open, configuration-driven platform that turns heterogeneous embodied-policy components into a reproducible data-to-deployment workflow. Rather than introducing another policy model, $\mathrm{FluxVLA}$ standardizes interfaces for datasets, visual-language and world models, action heads, reward- or advantage-weighted learning, distributed training, simulation evaluation, optimized inference, and robot operators. The engine further integrates compositional dual-arm simulation, scalable automatic data generation, and model-decoupled human-in-the-loop rollout, takeover, correction collection, and reward annotation. For responsive physical execution, it combines Real-Time Chunking (RTC) with accelerated inference backends, lightweight remote GPU serving, and configurable trajectory post-processing. Together, these capabilities connect offline learning, simulation validation, online correction, and real-robot execution through shared and auditable contracts. $\mathrm{FluxVLA}$ therefore targets the engineering bottlenecks separating promising embodied-learning algorithms from reproducible evaluation and dependable deployment. Code is available at https://github.com/FluxVLA/FluxVLA

发表机构

  • LimX Dynamics
  • Nankai University(南开大学)
  • The University of Hong Kong(香港大学)
  • University of Electronic Science and Technology of China(电子科技大学)
  • Beijing Institute of Technology(北京理工大学)
  • Xidian University(西安电子科技大学)
  • Zhejiang University(浙江大学)
  • Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑