发表机构
Microsoft(微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SharedKV-BT,利用行为树节点局部字段和并行评分,实现比自回归解码快2.36-4.15倍的决策,并提升操作任务准确率至94%、闭环成功率至60%。
AI 中文摘要
智能体任务需要一系列相互依赖的决策。自回归模型比传统分类器支持更灵活的决策接口,但会产生逐token生成的延迟。最近的共享前缀方法通过重用编码上下文并并行评分多个决策来降低这一成本,但不建模决策依赖关系或验证执行。我们提出SharedKV-BT,其中行为树(BT)的每个活动节点暴露阶段局部字段和候选,Shared-KV并行评分候选并将选定决策传递给单独的执行系统。我们在机器人操作、移动导航和计算机使用任务上测试了SharedKV-BT。在三个任务中,SharedKV-BT的类型化决策比提示匹配的自回归解码快2.36-4.15倍。在操作任务上,节点局部Shared-KV将联合决策准确率从75%提高到94%,闭环成功率从0%提高到60%。固定评分策略回放表明,阶段门控防止了乱序动作,外部后置条件防止了过早完成。
英文摘要
Agent tasks require sequences of interdependent decisions. Autoregressive models support more flexible decision interfaces than conventional classifiers but incur the latency of token-by-token generation. Recent shared-prefix methods reduce this cost by reusing encoded context and scoring multiple decisions in parallel, but do not model decision dependencies or verify execution. We propose SharedKV-BT, where each active node of a behavior tree (BT) exposes stage-local fields and candidates, and Shared-KV scores the candidates in parallel and passes the selected decision to a separate execution system. We tested SharedKV-BT on robot manipulation, mobile navigation, and computer-use tasks. Across three tasks, SharedKV-BT made typed decisions 2.36-4.15 times faster than prompt-matched autoregressive decoding. On the manipulation task, node-local Shared-KV improved joint decision accuracy from 75% to 94% and closed-loop success from 0% to 60%. Fixed-score policy replay showed that stage gating prevented out-of-order actions and external postconditions prevented premature completion.
Comments8 pages, 5 figures, 1 table. Last updated on October 5th, 2026