arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当重规划成为瓶颈:具身智能体的预算型重规划

When Replanning Becomes the Bottleneck: Budgeted Replanning for Embodied Agents

Shuaijun Liu, Feiyang You, Xingwei Chen, Ningxin Su

arXiv 2608.01428首次发表:更新:

AI 中文总结

针对具身智能体重规划延迟超标的瓶颈问题,提出搭载E-RECAP渐进式token剪枝方法的BRACE控制器,通过预算控制循环减少token消耗、降低SLO违反率,在多平台任务中提升了重规划的实时性与成功率。

AI 中文摘要

具身智能体频繁重规划以应对执行偏差、部分可观测性和协调风险,但每次基于大语言模型(LLM)的重规划调用会消耗随时间和智能体数量增长的累积文本上下文。一旦上下文规模过大,重规划延迟会出现重尾分布,即使任务成功率很高,也可能错过实时截止期限,这种失败模式仅通过平均延迟或成功率难以检测。我们提出BRACE,这是一个控制器,它将重规划构建为预算控制循环,通过决定是否重规划、选择重规划模式、分配显式的token预算和延迟服务水平目标(SLO),同时考虑可选的效率模块。作为可复用组件,我们引入E-RECAP,一种成本感知的渐进式token剪枝方法,它预测token效用并在Transformer层之间剪枝重规划上下文,同时保留关键的首尾token。在Meta Habitat、RoboFactory和AirSim平台上,当任务成功率已饱和时,搭载E-RECAP的BRACE将重规划调用的token数量减少了62%-92%,SLO违反率从85.5%-100.0%降至4.7%-50.0%。在更具挑战性的RoboFactory设置中,开环、冻结规划和无BRACE的方法均失败,而BRACE+E-RECAP达到80.0%的成功率,SLO违反率为4.6%,证明了感知重尾的逐调用预算控制在各类具身平台上的有效性。

英文摘要

Embodied agents replan frequently to recover from execution drift, partial observability, and coordination hazards, but each LLM-based replanning call can consume an accumulated textual context that grows over time and across agents. Once this context becomes large, replanning latency develops heavy tails and can miss real-time deadlines even when task success remains high, a failure mode that is hard to detect from average latency or success alone. We present BRACE, a controller that formulates replanning as a budgeted control loop by deciding whether to replan, selecting a replanning mode, and allocating an explicit token budget and latency service-level objective (SLO) while accounting for optional efficiency modules. As a reusable component, we introduce E-RECAP, a cost-aware progressive token pruning method that predicts token utility and prunes replanning contexts across transformer layers while preserving critical head and tail tokens. Across Meta Habitat, RoboFactory, and AirSim, BRACE with E-RECAP reduces replanning-call token counts by 62-92% and SLO violation rates from 85.5-100.0% to 4.7-50.0% in settings where task success is already saturated. In a harder RoboFactory setting where open-loop, frozen-plan, and No BRACE all fail, BRACE + E-RECAP reaches 80.0% success with 4.6% SLO violations, demonstrating that tail-aware per-call budgeting is effective across embodied platforms.

Comments20 pages total: 9 pages main text, 3 pages references, and 8 pages appendix; 18 figures and 32 tables. Accepted at ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑