arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Second Thought:LLM智能体行动与观察时的并行推理

Second Thought: Reasoning in Parallel as LLM Agents Act and Observe

Zhensu Sun, Chengran Yang, Yunbo Lyu, Jieke Shi, David Lo

arXiv 2608.13667首次发表:更新:

AI 中文总结

本文提出无需训练的推理框架Second Thought,利用LLM智能体行动与观察的空闲窗口并行推理,在多基准测试中降低轮次、减少主线程解码量,且Pass@1表现优于计算匹配对照组。

AI 中文摘要

ReAct范式中的LLM智能体会交替进行推理、行动与观察,但审慎推理仅局限于Thought阶段:智能体序列化一个行动并等待环境反馈时,其推理处于冻结状态。我们将行动与观察的这段重复区间识别为推理空闲窗口,探究其是否可承载服务于后续轮次的额外并行推理。为此,我们提出Second Thought,这是一个无需训练的推理框架,它在每个Thought阶段结束时立即分叉出四个辅助分支,与主循环并行解码,待环境观察结果到达时将生成的思考合并回去。通过这种方式,Second Thought将额外推理从主线程的顺序解码路径中移出。在三个智能体基准测试和三个推理LLM上,Second Thought在全部9组(模型、基准)对中降低了平均轮次,在其中6组中减少了主线程解码量,最高达43%(这些设置中平均约20%),而在第7组中基本保持不变;Pass@1在9组中的7组无显著变化,两个显著差异分别为+12.4和+10.2个百分点。与将同等计算预算强制分配给主线程自身推理的计算匹配对照组相比,在对照组适用的全部4组设置中,Second Thought均取得严格更高的Pass@1,且顺序解码量减少1.3至3.2。

英文摘要

LLM agents in the ReAct paradigm alternate between reasoning, acting, and observing, but deliberate reasoning is confined to the Thought phase: while the agent serializes an action and waits for the environment, its reasoning is frozen. We identify this recurring interval for Action and Observation as a reasoning idle window and ask whether it can host additional reasoning in parallel that serves future turns. Therefore, we propose Second Thought, a training-free inference framework that forks four auxiliary branches the instant each Thought phase concludes, decodes them concurrently with the main loop, and merges the generated thoughts back when the environment observation arrives. In this way, Second Thought relocates the added reasoning off the main thread's sequential decoding path. Across three agentic benchmarks and three reasoning LLMs, Second Thought lowers the average turn count in all nine (model,benchmark) pairs and reduces main thread decoding in six of them by up to 43% (roughly 20% on average among those settings), while leaving it essentially unchanged in a seventh; Pass@1 shows no significant change in seven of nine pairs and the two significant differences are +12.4 and +10.2 points. Against a compute-matched control that forces an equivalent budget onto the main thread's own reasoning, it attains strictly higher Pass@1 with 1.3 to 3.2 less sequential decoding in all four settings where the control applies.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑