基于常识的抽象指令路径规划
Commonsense-Grounded Path Planning from Abstract Instructions
浏览论文内容
中文总结 AI 辅助
CoRS利用大语言模型和视觉语言模型,将抽象指令转化为遵循常识的路径,通过常识排序发现未言明的考量并规划绕行,实验验证其优于现有LLM规划器。
中文摘要 AI 辅助
我们提出了常识排序搜索(CoRS),这是一种新颖的路径规划器,能将抽象指令转化为遵循常识的路线。现有方法尊重预先写下的考量,但与人共处的机器人也必须遵循那些未言明的考量,例如工人无需告知就会避开湿滑地面。CoRS利用大型语言模型(LLMs)和视觉语言模型(VLMs)作为常识知识,在规划中推理这些潜在考量。给定抽象指令(如“小心移动”)和环境中每个区域的视觉观察,CoRS为每个区域推导出考量,如“这个湿滑地面很滑,值得绕行”。然后,它比较区域之间的考量,以判断机器人应更避开哪两个区域,如“人群比湿滑地面更糟糕”。这些判断将区域排序为常识排名,其成本驱动传统搜索,始终返回有效路线。我们构建了一个潜在考量下的规划基准,包含三个环境、1350个问题和五个不同抽象级别的指令。实验表明,CoRS能发现未言明的考量,并绕过值得绕行的区域,同时穿过其余区域,这是近期基于LLM的规划器未能实现的行为。
英文摘要
We present \emph{commonsense ranked search} (CoRS), a novel path planner that turns an abstract instruction into a route that follows commonsense. While existing methods respect the considerations written down in advance, a robot working among people must follow those left unstated too, as with a wet floor that a worker avoids without being told. CoRS leverages large language models (LLMs) and vision-language models (VLMs) as commonsense knowledge to reason about these latent considerations in its planning. Given an abstract instruction (\emph{e.g.}, ``move carefully'') and visual observations of each region in the environment, CoRS derives a consideration for each region, as in ``this wet floor is slippery and worth a detour.'' It then compares the considerations between regions to see which of the two the robot should avoid more, as in ``the crowd is worse than the wet floor.'' These judgments sort the regions into a commonsense ranking, whose costs drive a conventional search that always returns a valid route. We build a benchmark for planning under latent considerations, with three environments, 1350 problems, and five instructions at three levels of abstraction. Experiments show that CoRS discovers the unstated considerations and goes around the ones worth a detour while crossing the rest, a behavior that recent LLM-based planners do not achieve.
发表机构
- CyberAgent Inc.(CyberAgent公司)
- Nagoya University(名古屋大学)
机构由 AI 辅助整理,请以论文原文为准。