arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

拆分何时划算?智能体大语言模型推理的预填-解码-注意力-前馈网络专业化模拟

When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference

Przemyslaw Forys, Haoran Wu, Can Xiao, Jiayi Nie, Tony Liu, Rika Antonova, Timothy Jones, Robert Mullins, Wayne Luk, Aaron Zhao, George A. Constantinides

arXiv 2608.03741首次发表:更新:

AI 中文总结

该研究提出HeteroPanacea模拟框架,针对智能体LLM推理的预填-解码-注意力-前馈网络拆分,验证其可提升吞吐量,为异构智能体服务系统提供支撑。

AI 中文摘要

智能体推理现已主导大语言模型(LLM)推理领域,要求LLM主动参与具备工具调用能力的多轮交互。这给底层推理系统带来更复杂的工作负载:预填(prefill)和解码(decode)等服务阶段呈现出截然不同的行为,对计算和内存带宽能力有不同需求。因此,单一同质GPU系统难以支撑智能体推理,推动行业转向具备拆分服务能力的异构系统,例如新兴的Vera-Rubin平台(集成GPU与Groq LPU)。然而,异构系统各组件的最优硬件形态这一问题仍未得到充分探索。为此,我们提出一种名为\textbf{HeteroPanacea}的新型拆分服务模拟框架,该框架支持跨三个维度的系统级模拟:1)拆分量化;2)设备内与设备间的自动并行调度;3)PDAF(预填-解码-注意力-前馈网络)NPU架构异构性。结合这三个维度,我们为未来异构智能体服务系统提供了跨栈模拟框架。我们验证了预填-解码拆分的优势:与当前GPU的传统服务相比,模拟显示服务吞吐量提升高达75%;且在假设存在定制NPU的情况下,4路预填-解码-注意力-前馈网络拆分是提升不同模型吞吐量最稳定的方案。我们还通过一系列消融研究,探究了模型架构与拆分增益之间的关系。

英文摘要

Agentic inference now dominates the LLM inference landscape, requiring LLMs to actively engage in multi-turn interactions with tool-calling capabilities. This introduces a more complex workload for the underlying inference system: serving stages such as prefill and decode exhibit substantially different behaviors and demand distinct compute and memory-bandwidth capabilities. As a result, a single homogeneous GPU system now struggles to support agentic inference, motivating an industry shift toward heterogeneous systems with disaggregated serving capabilities, such as the emerging Vera-Rubin platform with GPUs and Groq LPUs. However, the question of what the optimal hardware should look like for each component in a heterogeneous system remains underexplored. To this end, we propose a novel simulation framework for disaggregated serving, termed \textbf{HeteroPanacea}, that enables system-level simulation across three dimensions: 1) disaggregated quantization, 2) automated intra- and inter-device parallelization scheduling, and 3) PDAF (prefill-decode-attention-FFN) NPU architectural heterogeneity. By combining these three axes, we provide a cross-stack simulation framework for future heterogeneous agentic serving systems. We confirm the benefit of Prefill Decode disaggregation, simulating increased serving throughput by up to 75\% compared to traditional serving with current GPUs and demonstrate 4 way Prefill Decode Attention FFN disaggregation is the most consistent for increasing throughput across different models, assuming custom NPUs. We also investigate the relationship between model architecture and gain from disaggregation by running a set of ablation studies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑