arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KerColle:解锁视觉-语言-动作模型中的细粒度GPU并发性

KerColle: Unlocking Fine-Grained GPU Concurrency in Vision-Language-Action Models

Anna Li, Christina Giannoula, Nandita Vijaykumar

arXiv 2609.22335首次发表:更新:

发表机构

University of Toronto; Max Planck Institute for Software Systems (MPI-SWS)(多伦多大学; 马克斯·普朗克软件系统研究所(MPI-SWS))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

KerColle提出轻量级GPU调度框架,通过在线SM利用率与内核资源需求动态协同调度,缓解队头阻塞,实现VLA模型两阶段并发,提升吞吐量高达28%。

AI 中文摘要

视觉-语言-动作(VLA)模型已成为下一代机器人技术的基础模型。高VLA推理吞吐量对于满足机器人的控制率要求至关重要。VLA模型包含两个阶段,即视觉-语言模型(VLM)和动作头,这两个阶段可以解耦,并在独立的机器人请求之间异步并发执行。通过对四个最先进的VLA进行详细特征分析,我们观察到在VLA推理过程中GPU严重利用不足,因为GPU的协作线程数组(CTA)调度器无法完全重叠VLA执行的两个独立阶段。我们确定这种低效是由硬件线程块调度器中的队头阻塞引起的。我们证明,先前的调度框架并未解决VLA并发带来的挑战。首先,不同机器人请求之间的独立阶段各自包含众多内核,并且在任何给定时间,有不同组合的内核并行执行。这使得静态或提前调度策略在很大程度上无效。其次,许多动作头算子都是短运行内核,并且此类内核数量众多。这为在线分析或基于抢占的机制留下了很小的空间。为了解决这些挑战,我们提出了KerColle,一个轻量级GPU调度框架,它利用在线流式多处理器(SM)利用率和单个内核资源需求,智能地动态协同调度内核,以通过(1)减轻队头阻塞和(2)协同调度具有互补资源需求的内核来有效重叠两个执行阶段。我们在模拟中,跨两种GPU架构,针对4个最先进的VLA模型,证明了KerColle可提供高达28%的吞吐量提升。

英文摘要

Vision-Language-Action (VLA) models have emerged as foundational models for next-generation robotics. High VLA inference throughput is critical for meeting the control-rate requirements of robots. VLA models comprise two phases, a vision-language model (VLM) and an action head, that can be decoupled and executed asynchronously and concurrently across independent robot requests. Through a detailed characterization of four state-of-the-art VLAs, we observe that GPUs are severely underutilized in VLA inference as the Cooperative Thread Array (CTA) scheduler of GPU is unable to fully overlap the two independent phases of VLA execution. We identify that this inefficiency is caused by head-of-line blocking in the hardware thread-block dispatcher. We demonstrate that prior scheduling frameworks do not address the challenges posed by VLA concurrency. First, the independent phases across different robot requests each comprise numerous kernels, and at any given time, there are different combinations of kernels that are executed in parallel. This makes static or ahead-of-time scheduling policies largely ineffective. Second, many of the action-head operators are short-running kernels and there are numerous such kernels. This leaves no headroom for online profiling or preemption-based mechanisms. To address these challenges, we present KerColle, a lightweight GPU scheduling framework that leverages online Streaming Multiprocessor (SM) utilization and individual kernel resource requirements to intelligently and dynamically co-schedule kernels to efficiently overlap the two phases of execution by (1) mitigating head-of-line blocking, and (2) co-scheduling kernels with complementary resource requirements. We demonstrate in simulation, across two GPU architectures, for 4 state-of-the-art VLA models, that KerColle delivers throughput gains of up to $28\%$.

Comments11 pages, 15 figures, PACT 2026

DOI:10.1145/3838684.3846872

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑