发表机构
Sungkyunkwan University(成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文探究冻结视觉编码器的NPU卸载对机器人策略训练的影响,构建异步训练流水线,实验发现该方法可降低训练能耗,但会增加训练时间、降低策略成功率。
AI 中文摘要
当机器人策略针对新任务或新数据集进行训练时,可冻结其视觉编码器,仅训练动作生成模块,以此降低训练成本。冻结操作会移除编码器的反向传播过程,但由于输入图像会发生变化,其前向传播仍需在每一步训练中运行,因此会持续消耗GPU算力。为此,本文探究将该计算任务转移至NPU这类低功耗AI加速器,是否能在考虑额外数据传输和更长训练时间的情况下降低总能耗,以及这对策略性能有何影响。本文构建了一条异步训练流水线,将GPU与NPU共同用于AR-Actor专用模块。其中,冻结的视觉编码器在Mobilint Aries2 NPU上以A8W8 INT8精度运行,而FP32动作专家则在NVIDIA GeForce RTX 5060 Ti GPU上进行训练。本文将仅使用GPU的基准设置与四种条件(L1至L4)进行对比,这四种条件逐步将NPU卸载范围从1个Transformer编码器层扩展至4个。每种设置均使用3个随机种子训练30000步。对于仅GPU设置,本文测量了GPU板卡功耗;对于NPU设置,则测量了GPU与NPU板卡的总功耗。结果显示,L1设置(卸载ResNet18和第1个编码器层)的单样本能耗降低了17.1%,L4设置(卸载ResNet18和全部4个编码器层)的单样本能耗降低了27.9%。相比之下,L1设置的单样本训练时间增加了15.2%,L4设置则增加了37.7%,峰值已分配GPU内存降低了19.8至20.7%。本文对15个训练得到的策略分别使用300个环境种子进行评估,总共完成了4500次模拟器回滚。仅GPU设置的综合成功率为93.33%,而NPU设置的综合成功率为91.44至92.89%。这些结果表明,冻结视觉编码器的NPU卸载可降低训练能耗,但与仅使用GPU的训练相比,会增加训练时间,并使策略成功率降低0.44至1.89个百分点。
英文摘要
When a robot policy is trained for a new task or dataset, its visual encoder can be frozen and only its action generation module trained, reducing training cost. Freezing removes the encoder's backward pass, but its forward pass must still run at every training step because the input images change, so it keeps consuming GPU compute. We therefore ask whether moving this computation to a low power AI accelerator such as an NPU can reduce total energy despite the added data transfer and longer training time, and how it affects policy performance. We built an asynchronous training pipeline that uses both a GPU and an NPU for the AR-Actor specialist. The frozen visual encoder runs in A8W8 INT8 on a Mobilint Aries2 NPU, while the FP32 action expert is trained on an NVIDIA GeForce RTX 5060 Ti GPU. We compared a GPU-only baseline with four conditions, L1 to L4, which gradually extend NPU offloading from one to four Transformer encoder layers. Each condition was trained for 30,000 steps with three random seeds. We measured GPU board power for the GPU-only condition and combined GPU and NPU board power for the NPU conditions. Energy per sample decreased by 17.1% in L1, which offloaded ResNet18 and the first encoder layer, and by 27.9% in L4, which offloaded ResNet18 and all four encoder layers. In contrast, training time per sample increased by 15.2% in L1 and 37.7% in L4, and peak allocated GPU memory decreased by 19.8 to 20.7%. The 15 resulting policies were each evaluated with the same 300 environment seeds, for a total of 4,500 simulator rollouts. The combined success rate was 93.33% for GPU-only and 91.44 to 92.89% for the NPU conditions. These results show that NPU offloading of a frozen visual encoder can reduce training energy, but it increases training time and lowers policy success rate by 0.44 to 1.89 percentage points compared with GPU-only training.
Comments6 pages, 4 figures, 4 tables