发表机构
University of Chicago; Stanford University; Together AI(芝加哥大学; 斯坦福大学; Together AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LEAP提出一种延迟框架指导的小型草拟模型训练方法,通过针对目标动作序列训练0.6B模型,使LLM智能体端到端加速高达60%,且支持在线训练,便于实际部署。
AI 中文摘要
LLM智能体在展开(rollout)过程中以速度缓慢著称。智能体逐步完成任务,每一步先进行推理,然后选择要执行的动作。在下一步和下一个动作开始之前,前一步必须完成。推测解码通过在推理阶段草拟并验证推理令牌来加速展开过程。近期的工作也开始在动作阶段应用类似的想法。这些工作使用现成的模型(通常较大)来为目标模型草拟动作提案以供验证。较大的草拟模型与目标模型匹配的频率更高,但提案耗时更长,而较小的现成模型速度快,但很少做出与目标模型相同的决策。我们提出了一个更一般的问题:是什么决定了动作推测的端到端加速?为了回答这个问题,我们为推测轮次开发了一个延迟框架。该框架比较一轮推测的收益与成本。收益取决于草拟模型对目标模型的预测准确度以及任务在结束前可执行的步骤数。成本来自草拟、等待目标验证以及执行工具。在该框架的指导下,我们引入了LEAP(学习高效动作提案),该方法通过针对目标动作序列进行训练,保持草拟模型小而精确。使用一个0.6B的小模型,LEAP在大多数决策上与目标模型一致,并使智能体的端到端墙钟时间最多加快60%,且任务成功率无系统性变化。在各种数据集、目标模型和草拟模型上,该框架解释了大部分测得的加速效果。我们还展示了草拟模型可以在没有先前轨迹收集的情况下进行在线训练,并达到与离线训练相当的性能,这使得LEAP在实际部署中具有实用性。
英文摘要
LLM agents are known to be slow in rollouts. An agent completes a task one step at a time. At each step, it reasons and then chooses an action to execute. The next step and action cannot start until the previous one has finished. Speculative decoding accelerates the rollouts at the reason phase by drafting and verifying the inference tokens. Recent works have also started to apply similar ideas at the action phase. These works use off-the-shelf models, usually large, to draft action proposals for target model to verify. Large drafters match the target more often but take longer to propose, while small off-the-shelf models are fast but rarely make the same decision as the target. We ask a more general question: what determines the end-to-end speedup of action speculation? To answer it, we develop a latency framework for the speculative round. The framework compares what a round gains with what it costs. The gain depends on how well the drafter predicts the target and on how many steps the task can take before it ends. The cost comes from drafting, from waiting for target verification and from executing tools. Guided by the framework, we introduce LEAP (Learning Efficient Action Proposals) which keeps the drafter small and makes it accurate by training it on the target actions sequences. With a small 0.6B model, LEAP agrees with the target on most decisions and makes agents up to 60% faster in end-to-end wall clock time, with no systematic change in task success. Across various datasets, target models and draft models, the framework accounts for most of the measured speedups. We also show the draft model can be online trained with no prior trace collection and match the performance of offline training, making LEAP practical to deploy in the real world.