AI 中文总结
研究跨层异构系统设计空间难探索的问题,提出应用驱动框架CHASE,通过分层图表示候选方案,用解耦两级循环避免联合硬件映射搜索难的问题,在稀疏计算和大语言模型工作负载上评估,实现加速并降本降耗。
AI 中文摘要
人工智能和高性能计算基础设施越来越多地服务于结合密集张量计算、稀疏内核、大内存占用和通信密集型集合的工作负载组合。支持这些组合需要在加速器、内存层、扩展架构和集群网络之间进行协调选择。由此产生的跨层异构系统(XHS)设计空间很难探索。我们提出了CHASE,一个应用驱动的框架,通过工作负载来搜索物理上可行的XHS架构。CHASE将候选方案表示为分层类型图,并拒绝违反部署约束的设计。它通过解耦的两级循环避免了棘手的联合硬件映射搜索。我们在稀疏计算和大语言模型工作负载上评估了CHASE。其映射器与穷举最优解的差距在6.06%以内,同时平均映射时间相对于PEFT减少了60.5%。计算模型误差平均为4.4-7.5%,通信验证再现了物理平台上的关键趋势。外部搜索在64次迭代内达到接近全局最优。端到端案例研究表明,稀疏工作负载有利于关键感知异构Pod,而大语言模型推理有利于扩展岛;由此产生的设计分别实现了6.20倍和2.12倍的几何平均加速,同时相对于基线降低了成本和功耗。
英文摘要
AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collectives. Supporting these portfolios requires coordinated choices across accelerators, memory tiers, scale-up fabrics, and cluster networks. The resulting Cross-layer Heterogeneous System (XHS) design space is difficult to explore: hardware choices change legal task mappings, while rack power, switch radix, cabling, and cost constraints invalidate many candidates. We present CHASE, an application-driven framework that searches physically feasible XHS architectures through the workloads they must execute. CHASE represents candidates as hierarchical typed graphs and rejects designs that violate deployment constraints. It avoids intractable joint hardware-mapping search with a decoupled two-level loop: an inner mapper translates hardware-independent workload DAGs into topology-aware event traces, a calibrated event-driven simulator evaluates each mapping, and an outer telemetry-guided optimizer evolves the hardware graph. We evaluate CHASE on sparse-computing and LLM workloads. Its mapper remains within 6.06% of exhaustive optima while reducing mapping time by 60.5% on average relative to PEFT. Compute-model errors average 4.4-7.5%, and communication validation reproduces key trends across physical platforms. The outer search reaches near-global optima within 64 iterations. End-to-end case studies show that sparse workloads favor criticality-aware heterogeneous pods, whereas LLM inference favors scale-up islands; the resulting designs deliver 6.20$\times$ and 2.12$\times$ geomean speedups, respectively, while reducing cost and power relative to the baselines.
Comments20 pages, 15 figures. The first five authors contributed equally. Zhenhua Zhu, Hongyang Jia, and Shuwen Deng are corresponding authors