AI 中文总结
研究旨在加速大语言模型智能体,提出为推测器配备三个在线记忆系统,通过从智能体过去轨迹学习提升预测质量。经六个基准测试,记忆增强推测在动作和观察预测上有显著提升,且无损并能跨模型推广。
AI 中文摘要
推测执行通过在环境空闲时使用较小、成本较低的模型来预测和预启动下一步,从而加速大语言模型智能体。然而,现有的推测器是无状态的,会丢弃任务之间的所有信息,阻碍预测质量随经验提升。我们为推测器配备了三个在线记忆系统,它们从智能体过去的轨迹中学习:一个跟踪动作序列统计的对比转移表、一个检索上下文相似片段的情景记忆和一个抑制重复错误的混淆跟踪器。我们在跨越三种推测类型(动作预测、观察预测和链式预测)的六个基准上评估了这种方法。记忆增强的推测在动作预测上相对准确率提高了19%-39%,在具有重复动作空间的观察预测任务上提高了2.5倍。随着记忆积累,这些收益持续增长,并能推广到不同成本的推测器模型。所有推测都是无损的,因为它在空闲时间运行,不会增加实际时间成本,且智能体的轨迹与非推测执行相同。
英文摘要
Speculative execution accelerates LLM agents by using a smaller, cheaper model to predict and pre-launch the next step while the environment is idle. However, existing speculators are stateless and discard all information between tasks, preventing prediction quality from improving with experience. We equip the speculator with three online memory systems that learn from past agent trajectories: a contrastive transition table tracking action-sequence statistics, an episodic memory retrieving contextually similar segments, and a confusion tracker suppressing recurring errors. We evaluate this approach on six benchmarks spanning three speculation types: action prediction, observation prediction, and chained prediction. Memory-augmented speculation yields a 19--39\% relative accuracy improvement on action prediction and up to a $2.5\times$ increase on observation prediction tasks with repetitive action spaces. These gains grow continuously as memory accumulates and generalize across speculator models of varying cost. All speculation is lossless because it runs during idle time at zero added wall-clock cost, and the actor's trajectory is identical to non-speculative execution.