发表机构
Sun Yat-sen University; China Mobile Internet Company Ltd.(中山大学; 中国移动互联网有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
EdgeAgent为边缘UMA架构上的多智能体LLM推理设计跨层系统,通过零拷贝张量并行与动态草稿预算及异步挂起机制,在Apple M4上实现1.77倍加速。
AI 中文摘要
新兴的多智能体LLM系统需要隐私保护的边缘部署,然而当前的推理系统难以应对这些协作工作流。具体而言,内存受限的解码阶段在统一内存架构(UMA)上引发严重的总线争用,使朴素的CPU-GPU协同执行陷入瘫痪。此外,多智能体工作负载中的投机解码面临起草难度极端波动的问题,在复杂推理与可预测的结构化生成之间交替变化。加之频繁的工具调用导致的停顿,这种高度碎片化的执行严重低估了硬件利用率,并破坏了传统的静态批处理。我们提出EdgeAgent,一个专为边缘UMA和多智能体工作负载协同设计的跨层推理系统。在微架构层面,它绕开僵化的图编译器约束,实现零拷贝且UMA感知的张量并行,利用非对称内存布局来充分饱和CPU和GPU计算单元。在调度层面,它根据实时序列可预测性动态分配草稿预算,以限制带宽浪费。同时,一种异步挂起与让出机制主动驱逐停滞的智能体,确保在不可预测的工具调用期间硬件持续饱和。在Apple M4 SoC上的广泛评估表明,仅UMA感知执行就比批处理投机解码带来1.29倍加速。加入智能体感知调度后,完整EdgeAgent系统在极端工具使用延迟下达到1.77倍加速。
英文摘要
Emerging multi-agent LLMs demand privacy-preserving edge deployment, yet current inference systems struggle with these collaborative workflows. Specifically, the memory-bound decode phase causes severe bus contention on unified memory architectures (UMA), paralyzing naive CPU-GPU co-execution. Furthermore, speculative decoding in multi-agent workloads faces extreme variance in drafting difficulty, alternating between complex reasoning and predictable structured generation. Compounded by frequent tool-induced stalls, this highly fragmented execution severely underutilizes hardware and defeats traditional static batching. We present EdgeAgent, a cross-layer inference system explicitly co-designed for edge UMA and multi-agent workloads. At the micro-architectural level, it bypasses rigid graph-compiler constraints to enable zero-copy UMA-aware tensor parallelism, utilizing asymmetric memory layouts to fully saturate both CPU and GPU compute units. At the scheduling level, it dynamically allocates draft budgets based on real-time sequence predictability to bound bandwidth waste. Concurrently, an asynchronous suspend-and-yield mechanism actively evicts stalled agents, ensuring continuous hardware saturation during unpredictable tool invocations. Extensive evaluations on an Apple M4 SoC demonstrate that the UMA-aware execution alone contributes a 1.29x speedup over batched speculative decoding. Adding the agent-aware scheduling lifts the full EdgeAgent system to a 1.77x speedup under extreme tool-use latencies.