arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Parason:揭示大语言模型推理中的子任务并行与试错并行

Parason: Revealing Subtask and Trial Parallelism in LLM Reasoning

Zhengyang Zhang, Zijian Zhang, Jiaxuan Gao, Shusheng Xu, Yi Wu, Song Han, Ligeng Zhu

arXiv 2608.24658首次发表:更新:

发表机构

Tsinghua University; NVIDIA; Massachusetts Institute of Technology(清华大学; 英伟达; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Parason揭示LLM推理的子任务与试错并行,用上下文无关文法转换轨迹,经PA-GRPO训练,在数学基准上实现约1.7倍加速且保持竞争力。

AI 中文摘要

测试时推理的扩展大幅提升了大语言模型(LLM)的问题解决能力,但标准自回归解码仍按顺序执行长推理轨迹,给困难任务带来严重延迟(最长可达数天甚至数周)。并行推理是自然的解决方案,但现有系统主要关注子任务并行,即模型将高级任务分解为可独立解决的小块,这种方法忽略了另一种普遍的并行形式:试错并行,即多个推测尝试并行探索、验证和聚合竞争假设。本文提出Parason,它能在LLM推理中揭示并学习这两种并行形式。我们的分析表明,试错并行占可并行推理计算的大部分(在DeepSeek-V4的HLE推理步骤中占65.5%),且在困难问题上愈发占主导地位。基于该分类,Parason通过上下文无关文法将顺序推理轨迹转换为结构化并行轨迹,再用并行感知分组相对策略优化(PA-GRPO)训练模型,其奖励函数同时平衡准确率、延迟和两种并行度。推理时,Parason通过工具调用执行学习到的并行结构,将理论上的节省转化为实际的墙钟加速。在AIME24和AIME25等数学推理基准上的实验显示,Parason实现了约1.7倍的平均加速,同时保持了有竞争力的准确率。

英文摘要

Scaling test-time reasoning has substantially improved the problem-solving ability of large language models (LLMs), but standard autoregressive decoding still executes long reasoning traces sequentially, creating severe latency for difficult tasks (up to days and weeks). Parallel reasoning offers a natural remedy. However, prior systems primarily focus on Subtask Parallelism, where the model learns to decompose a high-level task into smaller chunks that can be solved independently. This approach overlooks another pervasive form of parallelism: Trial Parallelism, where multiple speculative attempts explore, verify, and aggregate competing hypotheses in parallel. In this paper, we introduce Parason, which reveals and learns both forms of parallelism in LLM reasoning. Our analysis identifies Trial Parallelism as the majority of parallelizable reasoning computation (65.5% in DeepSeek-V4's reasoning steps in HLE), and it becomes increasingly dominant on hard problems. Guided by this taxonomy, Parason converts sequential reasoning traces into structured parallel trajectories with a context-free grammar, then trains models with Parallelism-Aware Group Relative Policy Optimization (PA-GRPO), whose reward jointly balances accuracy, latency, and the two parallelism ratios. At inference time, Parason executes the learned parallel structure through tool calls, translating theoretical savings to real-world wall-clock acceleration. Experiments on mathematical reasoning benchmarks including AIME24 and AIME25 show that Parason achieves an average acceleration about 1.7$\times$ while maintaining competitive accuracy.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑