arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Trident:统一PyTorch Triton工作负载的受保护分发与主机执行

Trident: Unifying Guarded Dispatch and Host Execution for PyTorch Triton Workloads

Jinjie Liu, Xiaoyan Liu, Shuhan Zhang, Wenjia Sun, Ruilin Yang, Chunlei Men, Yonghua Lin, Shaohua Li

arXiv 2609.37241首次发表:更新:

发表机构

The Chinese University of Hong Kong; Beijing Academy of Artificial Intelligence(香港中文大学; 北京人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Trident是一个编译器后端,通过特化缓存模块将守卫分发与主机执行统一编译,消除PyTorch Triton内核调用开销,实现最高1.47倍和1.68倍的端到端加速。

AI 中文摘要

用户编写的Triton内核能够在PyTorch中实现高性能GPU计算,但其端到端延迟可能仍由主机端编排主导,尤其是在设备执行时间较短的情况下。尽管此http URL可以为捕获的图生成原生主机包装器,但每次调用在到达包装器之前仍会经历运行时管理的特化查找、守卫评估和准备过程。我们提出了Trident,一个编译器后端,它消除了特化缓存命中路径上的这种重复开销。Trident引入了特化缓存模块(SCM),该模块将受保护的特化选择、参数和执行环境准备以及多个特化的主机执行编译到一个可执行模块中。一次调用进入SCM,当特化匹配时保持在编译代码中,仅当需要编译新特化时才返回Python。基于Torch-MLIR构建,Trident将守卫和主机端编排降低到原生代码,同时保留对受支持的ATen运算符的优化运行时实现的调用。我们在两个LLM上的评估显示,Trident在模型级端到端延迟上相比即时执行实现了最高1.47倍的加速,相比此http URL实现了最高1.68倍的加速。

英文摘要

User-written Triton kernels enable high-performance GPU computation within PyTorch, but their end-to-end latency can remain dominated by host-side orchestration, especially when device execution is short. Although torch.compile can generate native host wrappers for captured graphs, each invocation still passes through runtime-managed specialization lookup, guard evaluation, and preparation before reaching the wrapper. We present Trident, a compiler backend that removes this recurring overhead from the specialization cache-hit path. Trident introduces the Specialization Cache Module (SCM), which compiles guarded specialization selection, argument and execution-environment preparation, and host execution for multiple specializations into a single executable module. An invocation enters the SCM once, remains in compiled code when a specialization matches, and returns to Python only when a new specialization must be compiled. Built on Torch-MLIR, Trident lowers guards and host-side orchestration to native code while retaining calls to optimized runtime implementations of supported ATen operators. Our evalu- ation on two LLMs shows that Trident achieves up to a 1.47x speedup in model-level end-to-end latency over eager execution and up to 1.68x over torch.compile.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑