SonicSampler:用于语言模型采样和推测性验证的统一瓦片感知内核
SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification
浏览论文内容
中文总结 AI 辅助
研究针对语言模型推理采样中现有实现的局限,提出SonicSampler,它是统一的瓦片感知Triton内核套件,垂直融合采样流水线。核心方法是分层两阶段top-k算法,在异构推测性解码工作负载中比基线快达16倍,保持灵活批处理执行。
中文摘要 AI 辅助
在语言模型推理中的采样包括用于推测性解码的一组组合的对数几率处理、令牌选择和验证操作。然而,现有实现要么仅加速此流水线的子集,依赖多个内核启动,要么假设跨批次的均匀采样行为,限制了对动态服务工作负载的支持并阻碍了高效的CUDA图执行。我们提出了SonicSampler,这是一套统一的瓦片感知Triton内核,它将完整的采样流水线垂直融合到一个固定的、工作负载感知的执行模型中。我们的内核支持动态的每个请求采样行为,包括语法约束解码、重复、频率和出现惩罚、对数几率偏差、温度缩放、top-k/top-p/min-p过滤以及推测性验证——在单个批处理内核中同时保持完全与CUDA图兼容。我们方法的核心是一种新颖的分层两阶段top-k算法,与竞争基线相比实现了高达10倍的加速,并利用语言模型输出的低熵结构在大词汇表上实现高效选择。在异构推测性解码工作负载中,SonicSampler在保持灵活的批处理执行的同时,比最先进的基线实现了高达16倍的加速。
英文摘要
Sampling in LLM inference comprises a combinatorial set of logit processing, token selection, and verification operations for speculative decoding. However, existing implementations either accelerate only subsets of this pipeline, rely on multiple kernel launches, or assume homogeneous sampling behavior across a batch, limiting support for dynamic serving workloads and preventing efficient CUDA Graph execution. We present $\textbf{SonicSampler}$, a unified suite of tile-aware Triton kernels that vertically fuses the complete sampling pipeline into a fixed, workload-aware execution model. Our kernels support dynamic per-request sampling behaviors, including grammar-constrained decoding, repetition, frequency and presence penalties, logit bias, temperature scaling, top-$k$ / top-$p$ / min-$p$ filtering, and speculative verification - within a single batched kernel while remaining fully CUDA Graph-compatible. Central to our approach is a novel hierarchical two-stage top-$k$ algorithm that achieves up to $\textbf{10x speedup}$ over competitive baselines and exploits the low-entropy structure of LLM outputs to enable efficient selection over large vocabularies. Across heterogeneous speculative decoding workloads, SonicSampler achieves up to $\textbf{16x speedup}$ over state-of-the-art baselines while preserving flexible batched execution.
发表机构
- Together AI
机构由 AI 辅助整理,请以论文原文为准。