arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.12139cs.NI

面向智能网络的AI感知GPU原生报文处理技术

Toward Intelligent Networks via AI-Aware GPU-Native Packet Processing

Seyed Mohammad Mehdi Mirnajafizadeh, Yiwen Hu, Rhongho Jang

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出AGP框架,将GPU升级为主要数据路径控制器,实现GPU原生报文处理等功能,消除CPU在关键路径的影响,大幅降低延迟、提升能效与推理吞吐量。

中文摘要 AI 辅助

将内联人工智能(AI)模型集成到关键网络基础设施中,根本上受限于CPU介导的报文处理带来的高延迟和同步开销。传统GPU卸载方案及近期的CPU旁路框架虽试图缩小这一差距,但仍受限于专有生态系统,或仍依赖主机CPU及粗粒度批处理来协调有状态遥测与AI流水线执行。本文提出AGP这一新型框架,将GPU从被动加速器提升为主要数据路径控制器,探究当GPU自主掌控完整的报文到推理流水线时,架构会发生何种变化。通过实现GPU原生报文处理、无竞争的GPU内有状态聚合,以及用于持续AI推理的持久 mega-kernel,AGP将CPU排除在关键路径之外。评估表明,原生GPU内报文处理可维持线速吞吐量,同时能效提升超6.1倍;此外,集成入侵检测系统(IDS)压力测试消除了传统CPU-GPU同步开销,端到端延迟降低7.9倍(p99延迟最高降低35倍),且仅使用约2%的GPU线程容量,就使全系统推理吞吐量提升9.3倍。

英文摘要

The integration of inline Artificial Intelligence (AI) models into critical network infrastructure is fundamentally bottlenecked by the high latency and synchronization overhead of CPU-mediated packet processing. While legacy GPU offload and recent CPU-bypass frameworks attempt to bridge this gap, they remain trapped in proprietary ecosystems or still rely on the host CPU and coarse-grained batching to coordinate stateful telemetry and AI pipeline execution. In this paper, we propose AGP, a novel framework that promotes the GPU from a passive accelerator to a primary data-path controller, answering what changes architecturally when the GPU autonomously owns the complete packet-to-inference pipeline. By enabling GPU-native packet processing, contention-free in-GPU stateful aggregation, and a persistent mega-kernel for continuous AI inference, AGP removes the CPU from the critical path. Our evaluation demonstrates that native in-GPU packet processing sustains line-rate throughput while being over 6.1x more power-efficient. Furthermore, our integrated Intrusion Detection System (IDS) stress test eliminates the legacy CPU-GPU synchronization tax, reducing end-to-end latency by 7.9x (up to 35x p99) and accelerating whole-system inference throughput by 9.3x using only ~2% of the GPU's thread capacity.

发表机构

  • Wayne State University(韦恩州立大学)
  • University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

机构由 AI 辅助整理,请以论文原文为准。

↑