EMO:人工智能工作负载的能源效率建模与优化
EMO: Energy Efficiency Modeling and Optimization for AI Workloads
浏览论文内容
中文总结 AI 辅助
针对GPU加速AI工作负载能耗大问题,EMO框架通过构建依赖图和假设分析确定优化位置,引入依赖感知内核打包确定优化方式,结合两者将能源优化转为约束组合问题求解,能在低开销下有效降低能耗。
中文摘要 AI 辅助
GPU加速的人工智能工作负载的巨大能耗对可持续计算构成挑战。我们发现执行异步(如CPU - GPU、并发流、多GPU)会产生空闲时间,使非关键内核能以更低频率运行以节省能源且不影响端到端延迟。然而,现有方法无法同时实现工作负载通用性和细粒度空闲时间发现,高保真建模开销大。我们提出EMO,一个利用这些细粒度机会的轻量级框架。首先,通过构建捕获异步的低级依赖图并进行假设分析来确定优化位置;其次,引入依赖感知内核打包以确定优化方式;最后,结合图分析和打包级模型将能源优化表述为约束组合问题并求解。评估表明EMO在仅2% - 5%性能损失和可忽略开销下将能耗降低15% - 28%。
英文摘要
The massive energy consumption of GPU-accelerated AI workloads challenges sustainable computing. We observe that execution asynchrony (e.g., CPU-GPU, concurrent streams, multi-GPU) creates slack, allowing non-critical kernels to run at lower frequencies to save energy without impacting end-to-end latency. However, existing approaches fail to simultaneously achieve workload generality and fine-grained slack discovery, while high-fidelity modeling incurs prohibitive overhead. We present EMO, a lightweight framework exploiting these fine-grained opportunities. First, to identify where to optimize, EMO constructs a low-level dependency graph capturing asynchrony and performs what-if timing analysis to precisely identify slack windows. Second, to determine how to optimize, EMO introduces dependency-aware kernel packing. It aggregates kernels to preserve critical paths while collapsing redundant details, enabling high-fidelity latency-energy modeling with minimal profiling cost. Finally, EMO combines graph analysis and pack-level models to formulate energy optimization as a constrained combinatorial problem, efficiently solving for optimal frequency policies under given latency targets. Evaluations show EMO reduces energy consumption by 15%--28% with only 2%--5% performance loss and negligible overhead.