arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19395cs.ARcs.AIcs.LG

HYDRA:面向动态混合大语言模型(LLM)工作负载的异构小芯片设计空间探索框架

HYDRA: A Heterogeneous Chiplet DSE Framework for Serving Dynamic Hybrid LLM Workloads

  • University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
  • University of Ulsan(蔚山大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiahao Lin, Alish Kanani, Sangwan Lee, Jaehyun Park, Umit Ogras

AI总结:

本文提出HYDRA框架,针对异构小芯片系统上混合LLM服务的设计空间开展联合探索,实现了吞吐量与首token生成时间的显著提升,凸显架构与运行时策略协同设计的重要性。

AI中文摘要:

混合Transformer-Mamba大语言模型(LLM)提升了长上下文处理效率,但其异构计算与通信模式为高效硬件加速带来了挑战。基于小芯片(chiplet)的架构通过集成专用计算与存储单元提供了可扩展解决方案,但涵盖静态架构配置与动态运行时策略的设计空间过大,难以穷尽探索。为应对这一挑战,本文提出HYDRA——一种面向异构小芯片系统上混合LLM服务的综合设计空间探索框架,其联合探索小芯片组成、布局、芯片间带宽配置、动态批处理及运行时调度,整合了感知通信的布局、动态批处理、弹性任务调度,以及一种快速马尔可夫性能估计器,该估计器可捕捉多租户运行时动态以实现高效且准确的探索。在所有工作负载上,HYDRA平均实现1.55倍吞吐量提升、43.7%的首token生成时间降低,与现有最优基线相比,吞吐量增益最高可达2.3倍,这些结果表明架构与运行时策略的协同设计对异构小芯片系统上高效大规模LLM服务至关重要。

英文摘要:

Hybrid Transformer-Mamba large language models (LLMs) enhance long-context efficiency, but their heterogeneous computation and communication patterns complicate efficient hardware acceleration. Chiplet-based architectures offer a scalable solution by integrating specialized compute and memory units. However, the design space spanning static architectural configurations and dynamic runtime policies is prohibitively large to explore exhaustively. To address this challenge, we present HYDRA, a comprehensive design space exploration framework for hybrid LLM serving on heterogeneous chiplet systems. HYDRA jointly explores chiplet composition, placement, inter-chiplet bandwidth provisioning, dynamic batching, and runtime scheduling. It integrates communication-aware placement, dynamic batching, elastic task scheduling, and a fast Markov-based performance estimator that captures multi-tenant runtime dynamics for efficient and accurate exploration. Across all workloads, HYDRA delivers 1.55x the throughput and 43.7 percent lower time-to-first-token on average, with throughput gains reaching up to 2.3x compared to state-of-the-art baselines. These results highlight that co-designing architecture and runtime policies is critical for efficient large-scale LLM serving on heterogeneous chiplet systems.

补充信息

↑