arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

EdgeDAE:面向边缘FPGA-GPU系统上具有微型VLA的实时物理AI的扩散动作专家加速

EdgeDAE: Acceleration of Diffusion Action Experts for Real-Time Physical AI with Tiny VLAs on Edge FPGA-GPU Systems

Zhiheng Chen, Ye Qiao, Mohammad Abdullah Al Faruque, Sitao Huang

arXiv 2610.00311首次发表:更新:

发表机构

University of California, Irvine(加州大学尔湾分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

EdgeDAE通过FPGA-GPU异构系统,将扩散动作专家的参数存储于片上,消除内存瓶颈,实现实时物理AI推理,延迟降低52.5%,能效提升17倍。

AI 中文摘要

物理AI模型(如视觉-语言-动作(VLA)架构)通过大规模Transformer骨干和基于扩散的动作解码器实现通用机器人策略。虽然边缘GPU平台擅长并行化计算密集型的视觉Transformer工作负载,但它们在扩散动作专家(DAE)模块上表现出根本性限制:迭代去噪过程需要在多个步骤中反复从DRAM加载参数,导致内存受限的性能,使得GPU的庞大计算吞吐量未被充分利用。DAE的I/O密集型特性与GPU以计算为中心的架构之间的这种不匹配,激发了异构加速方法。本文提出了EdgeDAE,一种异构FPGA-GPU系统,根据计算特性策略性地划分工作负载。我们将感知密集型的视觉Transformer卸载到GPU,同时通过BRAM/URAM中的完整片上参数存储,在FPGA上加速DAE推理。该架构通过协同设计量化策略、定点算术和面向FPGA结构的硬件高效随机数生成,消除了内存瓶颈。与边缘GPU基线相比,EdgeDAE将Octo-Small和Octo-Base的端到端推理延迟分别降低了52.5%和38.7%,吞吐量提高了2.10倍;与消费级GPU(RTX 4090)相比,其能效提高了约17倍。

英文摘要

Physical AI models such as Vision-Language-Action (VLA) architectures enable generalist robotic policies through large-scale transformer backbones and diffusion-based action decoders. While edge GPU platforms excel at parallelizing the compute-intensive vision-transformer workloads, they exhibit fundamental limitations for the Diffusion Action Expert (DAE) module: the iterative denoising process requires repeated parameter loading from DRAM across multiple steps, resulting in memory-bound performance where the GPU's massive computational throughput remains underutilized. This mismatch between DAE's I/O-intensive characteristics and GPU's compute-centric architecture motivates a heterogeneous acceleration approach. This paper presents \textbf{EdgeDAE}, a heterogeneous FPGA-GPU system that strategically partitions workloads based on computational characteristics. We offload the perception-heavy vision-transformer to GPU while accelerating DAE inference on FPGA through complete on-chip parameter storage in BRAM/URAM. This architecture eliminates the memory bottleneck by co-designing quantization strategies, fixed-point arithmetic, and hardware-efficient random number generation for the FPGA fabric. Compared to an edge GPU baseline, EdgeDAE reduces end-to-end inference latency by 52.5\% for Octo-Small and 38.7\% for Octo-Base, with up to $2.10\times$ higher throughput; compared to a consumer GPU (RTX~4090), it achieves ${\sim}17\times$ higher energy efficiency.

Comments7 pages, 5 figures. Accepted to ASP-DAC 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑