arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FusionML:预填充而非解码——统一内存苹果芯片上CPU+GPU协同执行的机制与边界

FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon

Om Mohite

arXiv 2607.22785首次发表:更新:

AI 中文总结

研究苹果硅片统一内存系统下变压器推理加速问题,找出此前尝试失败原因,提出基于急切实现边界的\sys{}方法,为变压器预填充实现CPU+GPU行拆分,经评估加速预填充,刻画了该方法在解码、训练等方面的边界并发布相关内容。

AI 中文摘要

苹果硅片系统级芯片通过统一内存系统共享CPU、GPU和神经引擎,这引发了一个问题,即能否通过跨单元拆分单个算子来加速变压器推理。之前包括我们自己的尝试都失败了,或者产生了精度混淆的结果。我们找出了原因:当CPU流操作在一个评估图中消耗未实现的GPU结果时,MLX的惰性图调度器会对跨流工作进行序列化。因此,一个行拆分矩阵乘法在输入实现时运行速度快1.38倍,但在惰性图中比仅使用GPU慢0.66倍;而急切的实现边界可恢复并发性(1.34倍)。\sys{}基于此修复为变压器预填充实现了每层的、感知竞争的CPU+GPU行拆分。在五个芯片和三代苹果硅片上进行的社区复制评估中,这种拆分将类似Llama的解码器块预填充加速了1.15至1.38倍,在完整的32块深度时保持不变,并且在通过标准MLX-LM提供服务的真实Qwen2.5 - 7B检查点上,首次生成令牌的时间快1.18至1.25倍,输出的令牌相同且解码吞吐量不变。我们同样仔细地刻画了边界:解码无法从中受益,受共享带宽限制,协同执行不会增加收益;精度匹配的训练在所有五个芯片上损失0.86至0.97;ANE调度开销在层粒度上排除了它;并且在内存压力下,无回归运行时门会适得其反,此时探测替代模式会逐出活动模式的工作集。代码、原始结果和生成记录已发布。

英文摘要

Apple-Silicon SoCs share CPU, GPU, and Neural Engine over one unified memory system, raising the question of whether transformer inference can be accelerated by splitting single operators across units. Prior attempts, including our own, failed or produced precision-confounded wins. We identify the cause: MLX's lazy-graph scheduler \emph{serializes} cross-stream work whenever a CPU-stream operation consumes an unmaterialized GPU result inside one evaluation graph, so a row-split matmul that runs \x{1.38} faster with materialized inputs runs \x{0.66} slower than GPU-only inside a lazy graph; an eager materialization boundary restores concurrency (\x{1.34}). \sys{} implements a per-layer, contention-aware CPU+GPU row split for transformer prefill built on this fix. Evaluated across five chips and three Apple-Silicon generations, community-replicated, the split accelerates Llama-shaped decoder-block prefill by \x{1.15}--\x{1.38}, unchanged at full 32-block depth, and reaches \x{1.18}--\x{1.25} faster time-to-first-token on a real Qwen2.5-7B checkpoint served through stock MLX-LM, with token-identical outputs and unchanged decode throughput. We characterize the boundaries equally carefully: decode cannot benefit, bound by shared bandwidth co-execution does not add; precision-matched training loses \x{0.86}--\x{0.97} on all five chips; ANE dispatch overhead excludes it at layer granularity; and a no-regression runtime gate becomes self-defeating under memory pressure, where probing an alternative mode evicts the active mode's working set. Code, raw results, and generation transcripts are released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑