arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ActTune:面向能效视觉-语言-动作推理的动作感知精度与GPU工作点自适应

ActTune: Action-Aware Precision and GPU Operating-Point Adaptation for Energy-Efficient Vision-Language-Action Inference

Zou Qingyun, Bin Gao, Wenju Zhao, Weng-Fai Wong, Bingsheng He, Tulika Mitra

arXiv 2610.08444首次发表:更新:

发表机构

National University of Singapore; Huazhong University of Science and Technology(新加坡国立大学; 华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA机器人推理的GPU能耗问题,提出ActTune框架,通过动作感知的逐层精度分配与GPU工作点自适应,在保持成功率的同时将每成功任务能耗降低高达76.8%。

AI 中文摘要

视觉-语言-动作(VLA)策略反复调用推理以控制机器人,使得图形处理单元(GPU)能耗成为任务执行中的经常性成本。然而,减少每次推理调用的能耗未必能降低每个成功任务的能耗,因为数值误差可能增加失败率,或推理速度变慢会延长执行时间。因此,我们以每个成功任务的GPU能耗为目标,同时保持任务成功率并将推理延迟增幅控制在10%以内。我们的方法基于两个观察:量化敏感性在不同动作类别、模型层以及权重与激活之间有所不同;数值精度改变工作负载,从而转移有利的GPU工作点。我们提出ActTune,一个动作感知框架,将逐层精度分配与依赖工作负载的GPU工作点选择(在请求的频率-功率上限对之间)相结合。一个轻量级决策树直接从配置动作误差中学习其分裂和叶精度配置,然后在每次策略调用前选择精度。控制器预测下一个工作负载,并使用在延迟预算下标定的查找表异步应用所选的GPU工作点。共享的驻留量化权重库支持配置切换,无需权重重建或额外的策略评估。在LIBERO(终身机器人学习基准)上,ActTune相对于最先进方法将平均任务成功率提高了最多2.3%。相对于原始BF16实现,它提供了最高2.02倍的推理加速,并通过GPU工作点自适应,将每个成功任务的能耗降低了最多76.8%。

英文摘要

Vision-language-action (VLA) policies repeatedly invoke inference to control robots, making graphics processing unit (GPU) energy a recurring cost of task execution. Reducing energy per inference call, however, may not reduce energy per successful task if numerical errors increase failures or slower inference prolongs execution. We therefore target GPU energy per successful task while preserving task success and keeping the inference-latency increase within 10\%. Our approach builds on two observations: quantization sensitivity varies across action classes, model layers, and weights versus activations; and numerical precision changes the workload, shifting favorable GPU operating points. We introduce ActTune, an action-aware framework that connects layer-wise precision allocation with workload-dependent GPU operating-point selection over requested frequency--power-cap pairs. A lightweight decision tree learns its splits and leaf precision configurations directly from configuration action errors, then selects precision before each policy call. The controller forecasts the next workload and applies the selected GPU operating point asynchronously using a lookup table calibrated under a latency budget. A shared resident quantized weight bank enables configuration switching without weight reconstruction or additional policy evaluations. On LIBERO, a benchmark for lifelong robot learning, ActTune improves mean task success by up to 2.3\% relative to state of the art. Relative to the original BF16 implementations, it delivers up to $2.02\times$ faster inference and, with GPU operating-point adaptation, reduces energy per successful task by up to 76.8\%.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑