HaReCAP:面向递归大语言模型智能体的习惯性动作 grounding 方法
HaReCAP: Habitual-action Grounding for Recursive Large Language Model Agents
浏览论文内容
中文总结 AI 辅助
HaReCAP 是针对 ReCAP 的低侵入性扩展,通过编译叶子反射规则减少长视距具身任务中 LLM 的重复调用,在 Robotouille 和 ALFWorld 上显著降低了 token 消耗。
中文摘要 AI 辅助
长视距具身任务要求大语言模型(LLM)智能体迭代分解高层目标、根据环境反馈修订计划,并将叶子级子目标 grounding 为可执行的有效动作。ReCAP 等递归上下文管理方法通过多级任务分解和父节点优化提升规划稳定性,但仍会在叶子节点重复调用 LLM,将原子子任务 grounding 为精确有效动作。我们将这一最终 grounding 步骤称为「最后一英里 grounding 冗余」,其在长视距执行过程中会累积为大量 LLM 调用和 token 开销。为缓解该问题,我们提出 HaReCAP(Habitual-action Grounded ReCAP),这是一种针对 ReCAP 的低侵入性叶子 grounding 扩展方法。HaReCAP 从成功轨迹中提取频繁的叶子决策,并离线编译为可审计、可弃权(不执行)的单步叶子反射规则。运行时,仅当规则能在当前有效动作集中唯一确定合法动作时,才跳过叶子节点的 LLM 调用;否则回退至原始 ReCAP。该设计避免了将完整递归上下文重复传入 LLM 以进行常规叶子动作 grounding,同时保留了原始递归控制流。我们以 Qwen3.5-27B 为主模型,在 Robotouille 和 ALFWorld 上评估 HaReCAP。在 ReCAP 和 HaReCAP 均能解决的任务中,HaReCAP 在 Robotouille 同步任务、Robotouille 异步任务、ALFWorld 上分别减少了 14.67%、17.93%、20.08% 的 token 消耗。结果表明,HaReCAP 可作为 ReCAP 类递归上下文管理框架的低侵入性扩展,在常见成功轨迹上减少跨环境和模型的最后一英里 grounding 冗余。
英文摘要
Long-horizon embodied tasks require LLM agents to iteratively decompose high-level goals, revise plans in response to environmental feedback, and ground leaf-level subgoals into valid executable actions. Recursive context-management methods such as ReCAP improve planning stability through multi-level task decomposition and parent-node refinement, but still repeatedly invoke the LLM at leaf nodes to ground atomic subtasks into exact valid actions. We refer to this final grounding step as last-mile grounding redundancy, which accumulates into substantial LLM-call and token overhead during long-horizon execution. To mitigate this issue, we propose HaReCAP (Habitual-action Grounded ReCAP), a low-intrusion leaf grounding extension for ReCAP. HaReCAP extracts frequent leaf decisions from successful trajectories and compiles them offline into auditable and abstainable one-step leaf-reflex rules. At runtime, it skips the leaf LLM call only when a rule can uniquely determine a legal action in the current valid-action set; otherwise, it falls back to the original ReCAP. This design avoids repeatedly carrying the full recursive context into the LLM for routine leaf action grounding, while preserving the original recursive control flow. We evaluate HaReCAP on Robotouille and ALFWorld with Qwen3.5-27B as the main model. On tasks solved by both ReCAP and HaReCAP, HaReCAP reduces token consumption by 14.67%, 17.93%, and 20.08% on Robotouille synchronous, Robotouille asynchronous, and ALFWorld, respectively. The results show that HaReCAP can serve as a low-intrusion extension to ReCAP-style recursive context-management frameworks, reducing last-mile grounding redundancy across environments and models on commonly successful trajectories.
发表机构
- North China Institute of Computer System Engineering(华北计算机系统工程研究所)
- University of Science and Technology of China(中国科学技术大学)
- China Information Security Research Institute Co., Ltd.(中国信息安全研究院有限公司)
机构由 AI 辅助整理,请以论文原文为准。