arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29379cs.RO

用受约束的大语言模型(LLM)弥合语义与物理鸿沟,实现安全可信的机器人操纵

Bridging Semantics and Physics with Constrained LLMs for Safe and Trustworthy Robotic Manipulation

  • Imperial College London(帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Wenhao Hong, Lan Wei, Dandan Zhang

AI总结:

该研究针对语言引导机器人操纵的语言-动作鸿沟,通过类型化契约约束LLM决策,结合MoveIt流水线验证,在多类机器人任务上实现了优于脚本策略的安全可靠性能。

AI中文摘要:

在真实厨房环境中运行的语言引导机器人,不仅要生成看似正确的计划,还需在感知不完善的杂乱环境中安全执行该计划。大语言模型(LLM)可将指令分解为动作序列,但存在语言-动作鸿沟:计划在语言层面看似有效,却可能因运动学和碰撞约束而在物理上不可行。我们通过将推理-执行边界形式化为类型化契约来弥合该鸿沟。系统基于RGB-D观测,在显式的碰撞感知场景模型中定位感知到的物体,并通过模型上下文协议(MCP)定义的模式验证工具调用约束语言层面的决策,在畸形命令到达机器人前将其拒绝。每个验证后的调用确定性地基于MoveIt Task Constructor流水线,在“先验证再执行”步骤中,针对重建的规划场景评估候选运动,仅通过运动学和碰撞检查的轨迹会被发送至机器人。在物理UFactory 850机器人上,该方法在涉及液体、颗粒介质和离散固体的倾倒任务中,每个任务10次试验的成功率最高达80%;在抓取-放置任务中,使用相同的规划、协议和验证栈,成功率达90%。尽管脚本策略在最简单任务上的表现略优于我们的方法,但其在最难任务上的成功率降至10%,而我们的方法在该任务上的成功率为60%。

英文摘要:

A language-guided robot operating in a real kitchen must do more than produce a plan that appears correct. It must also execute that plan safely in cluttered environments under imperfect perception. Large language models (LLM) can decompose instructions into action sequences, yet a language-action gap remains: a plan may appear valid linguistically while being physically infeasible under kinematic and collision constraints. We bridge this gap by formalizing the reasoning-execution boundary as a typed contract. From RGB-D observations, the system grounds perceived objects in an explicit, collision-aware scene model and constrains language-level decisions through schema-validated tool calls defined by the Model Context Protocol (MCP), rejecting malformed commands before they reach the robot. Each validated call is deterministically grounded in a MoveIt Task Constructor pipeline, where candidate motions are evaluated against the reconstructed planning scene in a verify-then-act step. Only trajectories that pass both kinematic and collision checks are sent to the robot. On a physical UFactory 850, the method achieves up to 80% success across ten trials per task on pouring tasks involving liquids, granular media, and discrete solids. It achieves 90% success on a grasp-and-place task using the same planning, protocol, and verification stack. Although a scripted policy slightly outperforms our method on the easiest task, its success rate falls to 10% on the hardest, compared with 60% for our method.

补充信息

↑