8 GB预算内的双臂操作:入门级Jetson上的零拷贝感知与量化ACT
Bimanual Manipulation Within an 8 GB Budget: Zero-Copy Sensing and Quantized ACT on an Entry-Level Jetson
浏览论文内容
中文总结 AI 辅助
该研究在8 GB入门级Jetson上实现了双臂操作,采用零拷贝感知、量化ACT,发现ACT在相同演示下优于Diffusion Policy,INT8量化可大幅降低延迟且保留任务成功率。
中文摘要 AI 辅助
通过模仿学习训练的双臂操作策略通常在工作站或数据中心级GPU上评估,而在嵌入式硬件上部署的成本却未被充分研究。我们提出了一个完全运行在NVIDIA Jetson Orin Nano Super(8 GB,NVIDIA嵌入式产品线的入门级产品)上的双臂SO-101系统,仅使用桌面级GPU(RTX 3070)进行离线训练,并通过可变形豆袋的抓取放置任务进行评估。首先,我们构建了由NVMM缓冲区支持的GStreamer捕获管道,消除了三相机感知中的冗余主机-设备拷贝。与预期相反,常规路径符合内存预算且未丢帧;零拷贝感知所带来的收益是CPU空闲空间(峰值单核利用率从98.0%降至77.0%)和最坏情况延迟(从117.31 ms降至101.52 ms)。其次,我们在相同的演示数据上分别训练ACT和Diffusion Policy,各自采用其参考预算(ACT为10万梯度步,Diffusion Policy为20万梯度步)。ACT收敛为具备任务能力的策略(20次试验中19次成功),而Diffusion Policy即使在两倍步长下也未收敛为可用策略(10次试验中0次成功),我们将此归因于不同的收敛成本而非精度上限。第三,我们将ACT转换为TensorRT。FP16将平均推理延迟从114.02 ms降至17.93 ms(加速6.4倍),INT8降至12.65 ms(加速9.0倍),三种精度下的任务成功率均保持(20次中19次、18次、19次)。我们报告了两项此前未记录的ACT相关发现:TensorRT的通用INT8校准会量化ResNet18骨干网络,但不接受145个Transformer层中的任何一个,这解释了尽管延迟进一步提升28%,INT8相比FP16的大小减少却可忽略(0.9%);量化的必要性取决于ACT的动作分块配置,在n_action_steps=100时全精度下可行,但在逐步重预测的时间集成中不可行。
英文摘要
Bimanual manipulation policies trained with imitation learning are typically evaluated on workstation or datacenter-class GPUs, leaving the cost of deploying them on embedded hardware largely uncharacterized. We present a bimanual SO-101 system running entirely on an NVIDIA Jetson Orin Nano Super (8 GB), the entry-level tier of NVIDIA's embedded line, using a desktop GPU (RTX 3070) only for offline training, evaluated on pick-and-place of a deformable beanbag. First, we build a GStreamer capture pipeline backed by NVMM buffers that removes redundant host-device copies from three-camera sensing. Contrary to expectation, the conventional path fit the memory budget and dropped no frames; what zero-copy sensing recovers is CPU headroom (peak single-core utilization 98.0% to 77.0%) and worst-case latency (117.31 ms to 101.52 ms). Second, we train ACT and Diffusion Policy on identical demonstrations, each at its own reference budget (100k gradient steps for ACT, 200k for Diffusion Policy). ACT converges to a task-competent policy (19/20 trials) while Diffusion Policy does not converge to a usable one (0/10) even at twice the step count, which we attribute to differing convergence costs rather than an accuracy ceiling. Third, we convert ACT to TensorRT. FP16 reduces mean inference latency from 114.02 ms to 17.93 ms (6.4x) and INT8 to 12.65 ms (9.0x), with task success preserved at all three precisions (19/20, 18/20, 19/20). We report two findings not previously documented for ACT: TensorRT's general-purpose INT8 calibration quantizes the ResNet18 backbone but accepts zero of 145 transformer layers, explaining INT8's negligible size reduction over FP16 (0.9%) despite a further 28% latency gain; and the need for quantization is conditional on ACT's action-chunking configuration, feasible in full precision at n_action_steps = 100 but not at the per-step re-prediction temporal ensembling requires.