arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越LLM服务:面向具身AI系统设计的视觉-语言-动作工作负载特征化

Beyond LLM Serving: Characterizing Vision-Language-Action Workloads for Embodied AI System Design

Seonghun Jung, Sieun Moon, Jiyoung Jeong, Jimin Lee, Jaehyuk Huh

arXiv 2610.05062首次发表:更新:

发表机构

KAIST(韩国科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究特征化视觉-语言-动作模型在批大小为1的机器人控制场景中的运行时行为,揭示动作张量维度、平台平衡和GPU频率缩放对性能的影响,并发现闭环运行中推理与动作执行重叠产生精度-速度-能量权衡,为具身AI系统设计提供指导。

AI 中文摘要

视觉-语言-动作(VLA)模型将多模态观测转换为低级机器人动作。在机器人运行期间,每个控制周期设定一个推理截止时间,超时会导致机器人基于过时观测执行动作,降低任务成功率。满足该截止时间促使在设备端或附近边缘执行推理,此时单个机器人需要批大小为1的推理,这超出了LLM服务系统的设计点。尽管VLA架构结合了熟悉的视觉-语言、自回归和扩散风格组件,但它们在批大小为1的控制场景中的运行时行为尚未被特征化。我们通过单次推理剖析和43,200次闭环回合,在边缘GPU服务器和两个片上系统(SoCs)上对四个代表性VLA模型进行了特征化。动作张量的维度决定了某个阶段是内存受限还是计算受限,平台平衡可以转移该瓶颈,GPU频率缩放产生一个依赖平台的能量-延迟最佳点。在闭环运行中,推理与动作执行的重叠产生了精度-速度-能量权衡,且没有配置在部署SLO(服务级别目标)中帕累托占优。这些结果指导VLA模型架构、硬件和运行时策略的联合设计。

英文摘要

Vision-language-action (VLA) models translate multimodal observations into low-level robot actions. During robot operation, each control period sets an inference deadline, and overruns leave the robot acting on stale observations, reducing task success. Meeting this deadline motivates on-device or nearby edge execution, where a single robot requires batch-1 inference outside the design point of LLM serving systems. Although VLA architectures combine familiar vision-language, autoregressive, and diffusion-style components, their runtime behavior in this batch-1 control setting remains uncharacterized. We characterize four representative VLA models on an edge GPU server and two onboard SoCs, using single-inference profiling and 43,200 closed-loop episodes. Action tensor dimensionality determines whether a stage is memory- or compute-bound, platform balance can shift that bottleneck, and GPU frequency scaling yields a platform-dependent energy-latency sweet spot. In closed-loop operation, overlapping inference with action execution creates an accuracy-speed-energy tradeoff, and no configuration is Pareto-dominant across deployment SLOs. These results guide joint design of VLA model architectures, hardware, and runtime policies.

CommentsTo appear in the Proceedings of the 32nd ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS '27). 17 pages, 14 figures

DOI:10.1145/3845814.3850089

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑