arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19315cs.ROcs.AI

GAVEL:用于验证和高效长时程LLM任务规划的图世界模型

GAVEL: Graph World Models for Verified and Efficient Long-Horizon LLM Task Planning

Ruiyang Wang, Hao-Lun Hsu, Swarajh Mehta, Jiwoo Kim, Zhihao Dou, Miroslav Pajic

首次发表
浏览论文内容

中文总结 AI 辅助

GAVEL通过显式图世界模型验证和修复LLM长时程规划,提升单任务成功率至91.8%,多任务至92.6%,并减少5.4%旅行距离。

中文摘要 AI 辅助

大型语言模型(LLMs)为长时程机器人规划提供了灵活的接口,但生成的规划常常无法遵守具身约束、从规划错误中恢复,或在部分可观测性下有效推理。我们提出GAVEL,一个围绕显式图世界模型构建的框架,用于验证和修复长时程LLM规划。该图表示相关的对象关系、动作前置条件和效果,以及关于未观测对象位置的概率信念。该模型可以在执行前预测LLM生成动作的后果,检测违规,并修复那些修正直接源自世界模型的违规。该方法还将LLM重新规划仅保留给需要语义推理的错误。对于多任务指令,GAVEL推理可能的对象位置分布,以重新排序剩余子任务并最小化预期搜索成本。我们在BEHAVIOR-1K上评估GAVEL,涵盖100个单一长时程任务和500个多任务指令。使用Qwen3-8B,GAVEL将单任务成功率从41.2%提升至91.8%,多任务成功率从19.9%提升至92.6%。与静态变体相比,分布性信念推理还减少了约5.4%的旅行距离。这些改进表明,显式图世界模型可以显著提高跨紧凑型和前沿托管LLM能力的长时程具身规划的可靠性和效率。

英文摘要

Large language models (LLMs) provide a flexible interface for long-horizon robot planning, but generated plans often fail to respect embodiment constraints, recover from planning errors, or reason effectively under partial observability. We present GAVEL, a framework for verifying and repairing long-horizon LLM planning built around an explicit graph world model. The graph represents relevant object-relations, action pre-conditions and effects, and probabilistic beliefs over unobserved object locations. This model can predict the consequences of LLM-generated actions before execution, detect violations, and repair those whose corrections follow directly from the world model. This method also reserves LLM replanning solely for errors requiring semantic reasoning. For multi-task instructions, GAVEL reasons over distributions of possible object locations to reorder remaining subtasks and minimize expected search cost. We evaluate GAVEL on BEHAVIOR-1K across 100 single long-horizon tasks and 500 multi-task instructions. With Qwen3-8B, GAVEL improves single-task success from 41.2% to 91.8% and multi-task success from 19.9% to 92.6%. Distributional belief reasoning also reduces travel distance by approximately 5.4% compared with a static variant. These improvements show that an explicit graph world model harness can substantially improve the reliability and efficiency of long-horizon embodied planning across compact and frontier hosted LLM capabilities.

↑