HeteroReason:面向分解式推测推理的异构FPGA-GPU加速
HeteroReason: Heterogeneous FPGA-GPU Acceleration for Disaggregated Speculative Reasoning
浏览论文内容
中文总结 AI 辅助
针对大型推理模型推测推理中前向轨迹缺乏鲁棒性和GPU资源利用不足的问题,提出算法-硬件协同设计的异构FPGA-GPU推理范式HeteroReason,通过回溯增强和预填充-解码分解实现平均4.2%准确率提升和1.01-1.42倍加速。
中文摘要 AI 辅助
大型推理模型(LRMs)通过利用思维链(CoT)推理,在推理任务中取得了最先进的性能。为了实现快速执行速度,推测推理技术采用轻量级草稿模型进行候选令牌生成,随后使用过程奖励模型(PRMs)进行验证,并由强目标模型进行细化。本文指出现有的推测推理范式遵循严格的前向推理轨迹,缺乏鲁棒性,如果早期推理步骤不理想,可能导致严重的错误传播。此外,在同类GPU平台上执行这些不同的推理方案(包括顺序草稿生成和并行验证)可能导致严重的资源利用不足。为解决这一问题,我们提出了HeteroReason,一种算法-硬件协同设计的异构FPGA-GPU推理范式,专门针对LRM推测推理。在算法层面,我们引入了一种增强回溯的工作流程,使系统能够从低质量状态中恢复并探索替代推理轨迹,显著提高推理鲁棒性。在系统层面,草稿模型被卸载到FPGA上,而PRM和目标模型部署在GPU上。我们优化了专门的工作流程以实现预填充-解码分解,利用影子同步将GPU侧的细化与FPGA侧的令牌更新重叠,有效隐藏同步延迟。为了缓解固有的顺序约束,我们提出了一种提前推测和细化调度方案,将系统从顺序执行方案转变为并行流水线。实验评估显示,与同类GPU基线相比,平均准确率提高了4.2%,延迟加速比为1.01倍至1.42倍,能效提升为1.25倍至1.57倍。
英文摘要
Large Reasoning Models (LRMs) have achieved state-of-the-art performance in reasoning tasks by utilizing Chain-of-Thought (CoT) reasoning. To achieve fast execution speed, speculative reasoning techniques adopt a lightweight draft model for candidate token generation followed by process reward models (PRMs) for verification and a strong target model for refinements. This paper identifies that the existing speculative reasoning paradigm follows a strictly forward-only reasoning trajectory, which lacks robustness and can lead to severe error propagation if early reasoning steps are suboptimal. Furthermore, executing these disparate inference schemes, including sequential drafting and parallel verification on homogeneous GPU platforms, can lead to severe resource underutilization. To address this, we propose HeteroReason, an algorithm-hardware co-designed heterogeneous FPGA-GPU inference paradigm specifically tailored for LRM speculative reasoning. At the algorithmic level, we introduce a backtracking-enhanced workflow that enables the system to recover from low-quality states and explore alternative reasoning trajectories, significantly improving reasoning robustness. At the system level, the draft model is offloaded to the FPGA while deploying the PRM and target models on GPUs. A specialized workflow is optimized to achieve prefill-decode disaggregation, which exploits shadow synchronization to overlap GPU-side refinements with FPGA-side token updates to effectively hide synchronization latency. To mitigate inherent sequential constraints, we propose a step-ahead speculation and refinement scheduling scheme, transitioning the system from a sequential execution scheme to a parallel pipeline. Experimental evaluations show an average 4.2% accuracy improvement, with 1.01x-1.42x latency speedups and 1.25x-1.57x improvements in energy efficiency compared to homogeneous GPU baselines.
发表机构
- Imperial College London(帝国理工学院)
- Tsinghua University(清华大学)
- University of Bristol(布里斯托大学)
机构由 AI 辅助整理,请以论文原文为准。