机器形状与分层分块:基于数组数学的形式化,附分层形状占用中的一个开放问题
Machine Shape and Hierarchical Blocking: A Mathematics of Arrays Formalization, with an Open Problem in Hierarchical Shape Occupancy
AI总结:
本文针对稠密矩阵乘法分块大小在不同芯片上预测偏差的问题,基于数组数学(MoA)框架形式化提出分层分块与开放问题,扩展算子推导分块调度,测试假设并指出框架局限,尝试扩展框架预测吞吐量。
AI中文摘要:
一项配套实证研究发现,针对Apple M1 Pro校准的稠密矩阵乘法分块大小,对应两个缓存公式在另一款芯片的已知缓存大小上预测偏差严重。本文将该研究提出的问题进行形式化。我们扩展了数组数学(Mathematics of Arrays, MoA)框架的数组形状推导算子,新增一个可从机器形状(即缓存容量、带宽及占用率组成的有序序列)推导分层多级分块与预取调度的新算子。该算子可将目前测试的三台真实机器上的校准值作为特例恢复,将每台机器的未知量缩减为少量与层级相关的占用率。随后我们明确提出本文的核心开放问题(未声称解决):这些占用率是否可从共驻性、私有缓存层级数量、关联性、预取器行为等更基础属性推导而来,还是本质上属于各架构特有的常数。我们提出四个可证伪假设并在真实硬件上测试,结果不一。我们还明确指出该框架的两个局限:其一,要使其有明确定义需专用非虚拟化硬件访问;其二,它仅部分扩展至分布式内存网络,在跨节点实现分块时需单独选择通信算法,而该框架本身不做此选择。本文最后以首次尝试扩展框架直接预测吞吐量(而非仅分块大小)收尾:两个项可从规格说明书推导,一个需单次测量,另一个在三台机器上测试后尚未实现跨机器迁移。
英文摘要:
A companion empirical study found that dense matrix multiplication block sizes calibrated on Apple M1 Pro correspond to two cache-t formulas that mispredict badly on a dierent chip's known cache sizes. This paper formalizes the question that nding raises. We extend the Mathematics of Arrays (MoA) framework's array-shape derivation operator to a new operator that derives a hierarchical, multi-level blocking and prefetch schedule from a machine's shape: an ordered sequence of cache-level capacities, bandwidths, and occupancy fractions. This operator recovers the calibrated values on every one of three real machines tested to date as a special case, reducing each machine's unknowns to a small number of level-specic occupancy fractions. We then state precisely, without claiming to resolve, the paper's central open problem: whether those fractions are derivable from more primitive properties co-tenancy, private-cache-level count, associativity, prefetcher behavior or are fundamentally per-architecture constants. Four falsiable hypotheses are stated and tested against real hardware, with mixed results. We further state two limits of the framework explicitly: it requires dedicated, non-virtualized hardware access to be well-dened at all, and it extends only partway to a distributed-memory network, where realizing a tile across nodes requires a separate choice of communication algorithm the framework does not itself make. A rst, honest attempt at extending the framework toward predicting throughput directly, not just block size, closes the paper: two terms prove derivable from a specication sheet, one requires a single measurement, and one tested across three machines does not yet transfer between them.