发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出消息传递语言模型(MPLM),通过线程间直接通信和抢占机制,在Sudoku、3-SAT和长上下文问答任务中实现比串行CoT和并行FJ更高效的推理。
AI 中文摘要
虽然推理时扩展提高了大型语言模型(LLM)的推理能力,但生成长思维链(CoT)的需求成为计算瓶颈。因此,与CoT等顺序扩展方法不同,最近的并行扩展技术使用分叉和连接(FJ)原语将工作分配到多个LLM线程中。然而,在分叉-连接范式中,线程通常是短暂的,并且不相互点对点通信,这限制了可扩展性。为了解决这个问题,我们引入了消息传递语言模型(MPLM),这是一种LLM推理框架,其中线程通过轻量级的发送和接收原语直接通信。MPLM通过两个关键机制实现高效扩展:(1)减少通信成本,通过避免冗余上下文共享实现;(2)抢占,允许线程基于来自同行的部分信息提前终止。我们在3类任务上展示了MPLM的前景。首先,在数独谜题上,我们表明MPLM需要的上下文比串行CoT和并行FJ渐近更小。然后,我们微调单个模型来解决25x25谜题,这些谜题对标准CoT和FJ方法以及没有工具的前沿推理模型仍然具有挑战性。其次,在3-SAT谜题上,抢占能力允许终止无希望的分支,从而提高了效率。最后,我们表明适当提示的大型预训练模型遵循MPLM协议,在长上下文问答上相对于流行的分叉-连接方法取得了有竞争力的结果。
英文摘要
While inference-time scaling has improved the reasoning abilities of large language models (LLMs), the need to generate long chains-of-thought (CoTs) is a computational bottleneck. Thus, in contrast to sequential scaling methods like CoT, recent parallel scaling techniques instead use fork and join (FJ) primitives to divide work across multiple LLM threads. However, in the fork-join paradigm, threads are typically transient and do not communicate pointwise with one another which limits scalability. To tackle this, we introduce Message Passing Language Models (MPLMs), a framework for LLM reasoning in which threads communicate directly via lightweight send and receive primitives. MPLMs enable efficient scaling through two key mechanisms: (1) reduced communication costs, achieved by avoiding redundant context sharing, and (2) preemption, which allows threads to terminate early based on partial information from their peers. We demonstrate the promise of MPLMs on 3 classes of tasks. First, on Sudoku puzzles, we show that MPLMs require an asymptotically smaller context than both serial CoT and parallel FJ. We then fine-tune a single model to solve 25 x 25 puzzles that remain challenging for standard CoT and FJ approaches, as well as frontier reasoning models without tools. Second, on 3-SAT puzzles, the capability of preemption allows termination of unpromising branches, which results in improved efficiency. Finally, we show that appropriately prompted large pre-trained models follow the MPLM protocol, achieving competitive results on long-context question answering relative to popular fork-join approaches.
CommentsCOLM 2026 (Oral Spotlight)