arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.03197cs.DC

弥合预测差距:补全机器形状使预测的时间、功耗、能量与映射与真实硬件行为一致

Closing the Prediction Gap: Completing Machine Shape So That Predicted Time, Power, Energy, and Mapping Match What Real Hardware Does

Lenore Mullin, Gaetan Hains

首次发表
浏览论文内容

中文总结 AI 辅助

本文通过索引五个未测量字段并引入 reached(d) 布尔值,补全机器形状以弥合预测与真实硬件行为之间的差距,在五个设备上验证并实现了所有字段的实测化。

中文摘要 AI 辅助

一篇配套论文从每个设备的实测形状 rho_machine(d) 推导出异构多机构设备集合的加权、无通信分区,并在真实硬件上进行了验证。其实验表明 rho_machine 是不完整的:fp32 到 fp16 的加速比和功耗变化未被完全索引,且内存容量预测依赖于假设而非实测的预留开销。本文弥合了这一差距,在 psi 选择词汇表中索引五个未测量的字段:内存容量 $M_i^{cap}$、功耗 $P_i$、执行单元能力 $X_d$、编译器指令配置 $C_d$ 以及内存预留开销 $R_d$。我们引入了 reached(d),一个布尔值,用于指示编译后的代码是否分派到专用硬件,从而解释了仅凭硬件可用性无法预测的 NVIDIA fp16 功耗变化。MoA 独立于目标计算,因此编译器指令获得两部分可接受性标准:选择真实硬件量而不进行任意重排序,并且是确定性的,不可被编译器覆盖。在 Intel、AMD、NVIDIA、OpenMP、OpenACC 和 Open MPI 中,这仅允许缓存放置控制、确定性向量或平铺宽度以及占用率上限;诸如自动调优之类的启发式方法被排除并吸收到残差 epsilon$(k)_d$ 中。我们在五个设备上进行了验证:A100、H100、V100、MI100 和 Max 1550,在每个设备上测量或确定了所有五个字段。可达性被证明是真实但部分的,占用率限制是设备特定的,预留开销范围约为 0.004 至 0.030。被排除的 Triton 自动调优器的残差为负 9.8%,表明未经验证的启发式方法可能产生与假定惩罚相反的效果。这些结果填补了配套集合中每个设备的所有五个字段,将假设行为转换为实测的、目标特定的量。

英文摘要

A companion paper derives a weighted, communication free partition for heterogeneous, multi institution device ensembles from each device measured shape, rho_machine(d), validated on real hardware. Its experiments show rho_machine is incomplete: fp32 to fp16 speedup and power shifts are not fully indexed, and memory capacity predictions rely on an assumed, not measured, reservation overhead. This paper closes that gap, indexing five unmeasured fields in the psi selection vocabulary: memory capacity $M_i^{cap}$, power draw $P_i$, execution unit capability $X_d$, compiler directive configuration $C_d$, and memory reservation overhead $R_d$. We introduce reached(d), a Boolean for whether compiled code dispatches to specialized hardware, explaining the NVIDIA fp16 power shift that hardware availability alone cannot predict. MoA fixes a computation independently of its target, so compiler directives get a two part admissibility criterion: select a real hardware quantity without discretionary reordering, and be deterministic, not compiler overridable. Across Intel, AMD, NVIDIA, OpenMP, OpenACC, and Open MPI, this admits only cache placement controls, deterministic vector or tile widths, and occupancy caps; heuristics like autotuning are excluded and absorbed into a residual epsilon$(k)_d$. We validate on five devices: A100, H100, V100, MI100, and Max 1550, measuring or establishing all five fields on each. Reachability proves real but partial, occupancy limits are device specific, and reservation overhead ranges about 0.004 to 0.030. A residual of negative 9.8 percent for the excluded Triton autotuner shows an unvalidated heuristic can act opposite a presumed penalty. These results close all five fields for every device in the companion ensemble, converting assumed behavior into measured, target specific quantities.

发表机构

  • College of Nanotechnology, Science and Engineering University at Albany, SUNY(奥尔巴尼州立大学纳米技术科学与工程学院)
  • LACL, Université Paris-Est Créteil(拉克莱特实验室,巴黎东 Créteil 大学)

机构由 AI 辅助整理,请以论文原文为准。

↑