arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经符号具身智能体

Neurosymbolic Embodied Agents

Mohammad Albinhassan, Yuming Feng, Alessandra Russo, Pranava Madhyastha

arXiv 2608.16794首次发表:更新:

发表机构

Imperial College London; Johns Hopkins University; City, University of London(帝国理工学院; 约翰斯·霍普金斯大学; 伦敦城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出一种神经符号具身智能体,将长周期家庭任务分解为视觉探索与符号规划,在VirtualHome、ALFWorld基准测试中表现优异,约束与搜索结合可大幅提升任务解决率,且无需专门训练。

AI 中文摘要

语言与视觉-语言模型能生成合理的具身计划,但无法保证可执行性,因为其输出可能违反环境动态或作用于错误接地的实体。我们提出一种神经符号智能体,它将长周期家庭任务分解为任务导向的视觉探索与受约束的符号规划。第一阶段,视觉-语言模型与探索工具从自我中心观测及接地交互中获取与目标相关的谓词和实例绑定,生成符号初始状态。第二阶段,PDDL(规划域定义语言)转换模型将解码限制为可扩展适用动作的标记,随后蒙特卡洛树搜索使用与域无关的规划启发式评估可执行的后续步骤。生成的计划在转换模型下天然可执行,其向环境的迁移取决于视觉接地是否正确。在VirtualHome与ALFWorld基准测试中,4B至27B规模的开放模型在两个环境中的成功率均超过90%,而我们的最小智能体在各环境中均大幅优于27B规模的直接视觉策略。约束与搜索呈互补而非可替代关系:在ALFWorld中,二者单独使用时解决的任务均不足三分之一,而结合后解决超过95%的任务。该方法生成的标记数仅为扩展思维方法的几分之一,且模型可见图像远少于直接交互方法,剩余失败均源于状态获取而非计划生成,无需任何专门训练。

英文摘要

Language and vision-language models generate plausible embodied plans but do not guarantee executability, as their outputs can violate environment dynamics or act on incorrectly grounded entities. We present a neurosymbolic agent that factors long-horizon household tasks into task-directed visual exploration and constrained symbolic planning. In the first phase, a vision-language model and exploration harness acquire goal-relevant predicates and instance bindings from egocentric observations and grounded interactions, producing a symbolic initial state. In the second, a PDDL transition model restricts decoding to tokens that extend applicable actions. Monte Carlo tree search then evaluates executable continuations using a domain-independent planning heuristic. The resulting plans are executable by construction under the transition model, with transfer to the environment conditioned on correct visual grounding. On VirtualHome and ALFWorld, open 4B-27B models exceed 90% success in both environments, and our smallest agent substantially outperforms a 27B direct visual policy in each. Constraints and search prove complementary rather than interchangeable: in ALFWorld either alone solves under a third of tasks, whereas their combination solves over 95%. The method also uses several times fewer generated tokens than extended thinking and far fewer model-visible images than direct interaction, and residual failures localize to state acquisition rather than plan generation without any specialized training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑