当更快的VLA部署改变闭环行为:SmolVLA在PyTorch与ONNX变体上的任务成功率-延迟分析
When Faster VLA Deployment Changes Closed-Loop Behavior: Task Success-Latency Analysis of SmolVLA Across PyTorch and ONNX Variants
AI总结:
本研究分析SmolVLA在PyTorch与ONNX部署下的闭环任务成功率与延迟权衡,发现ONNX加速导致Spatial成功率下降,而增加上下文宽度可恢复性能,强调部署评估需综合报告延迟、工件与闭环成功率。
AI中文摘要:
视觉-语言-动作(VLA)部署可以在降低推理延迟的同时改变闭环任务行为。我们在LIBERO Spatial和Object(MuJoCo 3.3.2,LeRobot 0.6.1,seed 42)中,于RTX 2060(6 GB)上评估了HuggingFaceVLA/smolvla_libero,比较了PyTorch+AMP与ONNX Runtime CUDA执行提供程序(CUDA EP)。主要评估使用100个回合/套件;配对滚动使用300个回合/套件。PyTorch+AMP在1181 ms p99下达到70.0%/88.0%的Spatial/Object成功率。Requested-FP16和requested-INT8 ONNX将tether-inspect p99降低至601 ms和532 ms,而Spatial成功率降至41.0%和40.0%,Object保持89.0%。图审计显示这些工件是字节相同的FP32图,因此requested-INT8行不是操作符级INT8量化。静态语言宽度消融(16/24/32个token)产生41.0%、75.0%和71.0%的Spatial成功率;宽度24和32恢复大部分Spatial下降,而Object成功率和统一基准延迟大致保持稳定。宽度24的ONNX Spatial成功率与PyTorch+AMP基线相当,延迟约为其一半(Wilson区间重叠;双比例卡方p=0.53)。上下文宽度是该堆栈中的一个重要贡献因素;它并不能解释PyTorch与ONNX之间的每一个差异。部署评估应联合报告延迟、工件检查、接口约束和闭环成功率。代码:此HTTPS URL。
英文摘要:
Vision-language-action (VLA) deployment can reduce inference latency while changing closed-loop task behavior. We evaluate HuggingFaceVLA/smolvla_libero on an RTX 2060 (6 GB) in LIBERO Spatial and Object (MuJoCo 3.3.2, LeRobot 0.6.1, seed 42), comparing PyTorch+AMP with ONNX Runtime CUDA Execution Provider (CUDA EP). The main evaluation uses 100 episodes/suite; a paired rollout uses 300 episodes/suite. PyTorch+AMP reaches 70.0%/88.0% Spatial/Object success at 1181 ms p99. Requested-FP16 and requested-INT8 ONNX reduce tether-inspect p99 to 601 ms and 532 ms, while Spatial success falls to 41.0% and 40.0% and Object remains at 89.0%. A graph audit shows those artifacts are byte-identical FP32 graphs, so the requested-INT8 row is not operator-level INT8 quantization. A static language-width ablation (16/24/32 tokens) yields Spatial success of 41.0%, 75.0%, and 71.0%; widths 24 and 32 recover much of the Spatial drop while Object success and uniform-bench latency stay approximately stable. Width-24 ONNX Spatial success is comparable to the PyTorch+AMP baseline at roughly half the latency (Wilson intervals overlap; two-proportion chi-squared p=0.53). Context width is an important contributor in this stack; it does not account for every PyTorch-vs-ONNX difference. Deployment evaluation should jointly report latency, artifact inspection, interface constraints, and closed-loop success. Code: https://github.com/rafiqul713/smolvla-libero-onnx.