自编排语言模型:利用语义依赖实现高效推理
Self-Orchestrating Language Models: Leveraging Semantic Dependence for Efficient Inference
- Massachusetts Institute of Technology(麻省理工学院)
- Google(谷歌)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出自编排语言模型,通过标注语义依赖来指导推理,实现并行解码、缓存驱逐和去噪排序,以帕累托最优方式平衡质量与效率。
AI中文摘要:
大型语言模型(LLMs)展现了令人印象深刻的能力,但其部署带来了显著的效率挑战。自回归解码在低批量大小场景下会造成严重的推理延迟,并且未能充分利用硬件加速器。离散扩散模型可以并行生成,但若没有大量扩散去噪步骤,则难以达到自回归质量。长上下文推理造成的内存瓶颈甚至使最先进的加速器也捉襟见肘。我的论点是,语言模型可以通过在其生成过程中标注语义依赖——即哪些词元依赖于哪些其他词元——来指导自身的推理执行策略。我将此类模型称为自编排语言模型。针对每个系统,我设计了一个运行时,该运行时根据这些标注来并行化自回归解码、驱逐中间上下文或推导去噪顺序,从而实现帕累托最优的质量-效率权衡。我通过三个自编排系统来展示这种方法。首先,PASTA利用语义依赖来并行化自回归解码,训练模型标注哪些输出块可以独立生成。其次,TIP利用语义依赖从KV缓存中驱逐中间推理步骤,在保持准确性的同时减少内存消耗。第三,Planned Diffusion利用语义依赖为离散扩散推导去噪顺序,自回归地生成一个计划,指定哪些块可以并行去噪。
英文摘要:
Large language models (LLMs) demonstrate impressive capabilities, but their deployment presents significant efficiency challenges. Autoregressive decoding imposes substantial inference latency and under-utilizes hardware accelerators in low batch size regimes. Discrete diffusion models can generate in parallel but struggle to match autoregressive quality without many diffusion denoising steps. Long-context reasoning creates memory bottlenecks that strain even state-of-the-art accelerators. My thesis is that language models can direct their own inference execution strategy by annotating semantic dependence -- which tokens depend on which others -- in their generation. I call such models self-orchestrating language models. For each system, I design a runtime that acts on these annotations to parallelize autoregressive decoding, evict intermediate context, or derive denoising orders, achieving Pareto-optimal quality-efficiency trade-offs. I demonstrate this approach through three self-orchestrating systems. First, PASTA uses semantic dependence to parallelize autoregressive decoding, training the model to annotate which output chunks can generate independently. Second, TIP uses semantic dependence to evict intermediate reasoning steps from the KV cache, reducing memory consumption while preserving accuracy. Third, Planned Diffusion uses semantic dependence to derive a denoising order for discrete diffusion, autoregressively generating a plan that specifies which chunks to denoise in parallel.