arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17717cs.RO

CompCPZ:在语言引导的机器人操作中保留多模态意图

CompCPZ: Preserving Multi-Modal Intent in Language-Guided Robot Manipulation

  • School of Computation, Information and Technology(计算、信息与技术学院)
  • Technical University of Munich(慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

Zhen Zhang, Ahmad Hafez, Peng Xie, Yanliang Huang, Wenyuan Wu, Amr Alanwar

中文总结 AI 辅助

CompCPZ是一种用于语言引导机器人操作的代数层,可恢复多模态析取表示,在ManiSkill3基准测试中性能优于多种基线,且能迁移至真实机器人实验,凸显组合语言接地需评估意图连通分量结构。

中文摘要 AI 辅助

当机器人被要求“将杯子放在红色盘子或蓝色盘子附近”时,它可能会到达两者之间的质心,在几何上看似成功,但却未满足指令的任何一个析取项。这种隐性语义缺陷暴露了语言条件机器人策略的结构局限:将析取指令压缩为单个连通集合的表示无法保留所有可行模态,而在运行时模态不确定性下,承诺单一动作的规划器性能会下降。我们通过CompCPZ解决这一局限,它是一种健全的代数层,语言条件学习系统可将其包裹以恢复多模态析取表示,沿语言解析树递归组合各基元的约束多项式zonotope包络,具备无分布保形覆盖能力,且运行时延迟低于1毫秒。在闭环ManiSkill3桌面操作基准测试中,CompCPZ优于凸集基线、多峰解码器及零样本视觉-语言-动作模型(1900/1918次配对获胜,p远小于10^(-30));同一编译器无需微调即可迁移至Unitree Go2四足机器人的平面真实机器人运动捕捉实验。这些结果表明,组合语言接地的评估不应仅看是否到达解码目标,还应看所表示的可行集合是否保留用户意图的连通分量结构。

英文摘要

A robot asked to "place the cup near the red plate or the blue plate" may reach the centroid between them and appear geometrically successful, while satisfying neither disjunct of the instruction. This silent semantic failure exposes a structural limitation of language-conditioned robot policies: representations that collapse a disjunctive instruction into a single connected set cannot preserve all feasible modes, and planners that commit to one action degrade under run-time mode uncertainty. We address this limitation with CompCPZ, a sound algebraic layer that language-conditioned learning systems wrap to recover multi-modal disjunctive representation, recursively composing per-primitive constrained polynomial zonotope enclosures along the language parse tree with distribution-free conformal coverage and sub-millisecond runtime. On a closed-loop ManiSkill3 tabletop-manipulation benchmark, CompCPZ outperforms convex set baselines, multi-peak decoders, and a zero-shot vision-language-action model (1,900/1,918 paired wins, p << 10^(-30)); the same compiler also transfers without retuning to planar real-robot trials on a Unitree Go2 quadruped under motion capture. These results suggest that compositional language grounding should be evaluated not only by reaching a decoded target, but by whether the represented feasibility set preserves the connected-component structure of the user's intent.

↑