arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SAGE:面向LLM任务规划器的符号动作门控与编辑

SAGE: Symbolic Action-Gating and Editing for LLM Task Planners

Trung Minh Bui, JongSul Moon, YoungOuk Kim, Quang-Ngoc Phung, Se-Woong Jun, Dongin Shin

arXiv 2609.34268首次发表:更新:

发表机构

Korea Electronics Technology Institute (KETI)(韩国电子技术研究院(KETI))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SAGE提出符号动作门控与局部编辑机制,以零token安全监控和高效恢复提升LLM任务规划器的安全性与完成率,并在多个基准上验证了其优势。

AI 中文摘要

大型语言模型(LLM)现已成为具身家庭智能体的默认认知核心,然而它们生成的计划在执行前很少依据环境的基础模型进行检查,而且它们报告的任务成功率往往是在高度饱和的基准上测得的,以至于任何方法都无法与其他方法区分开来。我们提出SAGE(符号动作门控与编辑),这是一个由两种轻量机制构建的单LLM规划器:一个与领域无关的符号门(约250行Python代码,零token,复杂度为O(|π|)),作为运行时安全监控器,以类型化原因阻止违反前提条件的动作;以及一个局部编辑机制,仅重新生成失败子目标的后缀,保留已完成和未受影响的工作;混合的种子加实时记忆存储支持冷启动覆盖。我们在无泄漏协议(留一法检索)下,对五个开放权重模型和一个75任务AI2-THOR基准进行了评估。在标准基准上,目标完成率趋于饱和(52%的实例被平凡解决),SAGE与强分层基线持平。在更困难、与方法无关的多目标组合上,SAGE的完成率优势重新显著出现(四个模型上+0.06至+0.23)。在注入的中途执行失败下,SAGE以2.4至3.3倍更少的LLM调用,与全计划重规划器一样可靠地恢复。作为执行前验证门,符号监控器在执行前阻止不安全动作,并为每个测试的规划器提高模拟器报告的分步成功率(最高+0.11),这是验证器从未见过的信号(非循环)。由于该门不调用任何模型(0.008毫秒/计划),它是一个在边缘上几乎免费运行的安全层:SAGE规划在Jetson AGX Orin上复现了其质量,其中小模型验证帮助最大。我们发布了基准、无泄漏协议、恢复和安全门工具集,以及验证器可移植性研究(在ALFWorld上自动诱导,保留0.89)。

英文摘要

Large language models (LLMs) are now the default cognitive core of embodied household agents, yet the plans they emit are rarely checked against a grounded model of the environment before execution, and the task-success they report is often measured on benchmarks so saturated that no method can be separated from another. We present SAGE (Symbolic Action-Gating and Editing), a single-LLM planner built from two lightweight mechanisms: a domain-agnostic symbolic gate (~250 lines of Python, zero tokens, $O(|π|)$) that blocks precondition-violating actions with typed reasons as a runtime safety monitor, and a local edit that regenerates only the failed sub-goal's suffix, keeping completed and untouched work intact; a hybrid seed+live memory store supports cold-start coverage. We evaluate under a leak-free protocol (leave-one-out retrieval) over five open-weight models and a 75-task AI2-THOR benchmark. On the standard benchmark goal-completeness saturates (52% of instances trivially solved) and SAGE ties strong hierarchical baselines. On a harder, method-agnostic multi-goal composition, SAGE's completeness lead re-emerges large (+0.06 to +0.23 across four models). Under injected mid-execution failures, SAGE recovers as reliably as whole-plan replanners at 2.4-3.3x fewer LLM calls. As a verify-before-execute gate, the symbolic monitor blocks unsafe actions before actuation and raises simulator-reported step-success for every planner tested (up to +0.11), a signal the verifier never sees (non-circular). Because the gate calls no model (0.008 ms/plan), it is a safety layer that runs essentially free on the edge: SAGE planning reproduces its quality on a Jetson AGX Orin, where small-model verification helps most. We release the benchmark, the leak-free protocol, the recovery and safety-gate harnesses, and a verifier-portability study (auto-induced on ALFWorld, 0.89 held-out).

Comments8 pages, 2 figures, 1 table, 3 algorithms. Submitted to IEEE Robotics and Automation Letters (RA-L). Code and benchmark: https://github.com/mtbui2010/sage_release

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑