ForgeStencil:从内核到100多个真实应用的逐案例模板自动特化
ForgeStencil: Automating Per-Case Stencil Specialization from Kernels to 100+ Real Applications
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
ForgeStencil利用代码合成代理实现逐案例模板特化,在内核和100多个真实应用中分别达到2.35倍和1.41倍的加速,超越通用方法。
AI中文摘要:
在现代GPU上,最快的模板(stencil)内核取决于模板的形状、精度和宿主应用程序,而针对一种情况调优的内核很少是另一种情况的最快选择。模板领域特定语言(DSL)、代码生成器和自动调优器反而追求通用性:一种由人类编写的方法跨情况复用,并主要在微基准测试上验证,因为逐案例特化成本过高而难以扩展。ForgeStencil从相反的假设出发。代码合成代理已将该成本降低到足以针对每种情况构建全新解决方案,并在真实软件中端到端部署。一个内核代理合成CUDA代码,并锻造一个针对每种配置的特化算子矩阵,其性能匹配或超过每种情况下公开可用的最强现有技术(SOTA)基线。一个应用代理将该原则扩展到整个应用程序:它定位热点,重写应用程序结构,并在100多个真实工业和科学代码中验证和集成每个更改。大部分测得的加速来自结构性和宿主端重写,纯模板替换占少数;增益还与基线已调优的程度呈负相关,这与增益来自特化而非通用复用一致。每个结果都通过测量完整性检查工具验证,该工具将夸大的加速转化为系统级错误。锻造的内核在相同精度f32下,相对于逐案例SOTA基线的几何平均加速为2.35倍(fp16增益为1.95倍,另行披露;A100),端到端应用程序中位数在100个代码上为1.41倍,所有对比均基于相同架构的GPU基线,并带有程序提供的验证和计时。在116个候选项中,每个未通过正确性、测量或加速标准的候选都被记录为拒绝或降级,而不是被写为加速。
英文摘要:
Industrial and scientific computing rests on a few core kernels, and the stencil is among the most widely used: weather and climate models, seismic imaging, fluid dynamics, and image processing all run on it. No single stencil implementation is fastest: the optimal kernel changes qualitatively with stencil shape, grid shape, precision, and host application. For two decades the field has answered with general methods (DSLs, code generators, autotuners), because specialized solutions were too expensive to build per case, so all reuse one human-authored recipe. That reuse costs performance; we call the cost the generality tax. This premise no longer holds: code-synthesis agents now build a correct, specialized solution per case at acceptable cost. ForgeStencil automates this. A Kernel Agent synthesizes CUDA and forges a per-configuration map of specialized operators, removing the tax case by case. On an A100 the map beats the strongest public baseline in 37 of 37 cases: geometric mean 2.35x against same-precision f32 baselines and 1.95x for fp16, each reported under its own precision. The same change reaches end-to-end application performance. A generic operator library is tuned once for its own general case and reused across applications, so its shapes, layouts, and launch boundaries are optimal for none of them: using it is the application-level form of the tax. An App Agent instead forges a specialized solution per application, locating hotspots, rewriting application structure, and validating and integrating each change. Across 100 real industrial and scientific codes the end-to-end median speedup is 1.41x against each application's own GPU baseline. To our knowledge this is the first demonstration that per-case synthesis carries from a kernel library to complete applications at this breadth, and evidence that reuse is no longer the default in a domain built on it for two decades.