缓解跨模态干扰并确保几何可行性:基于可供性引导与自洽多模态大语言模型的指令跟随操作任务规划
Mitigating Cross-Modal Distraction and Ensuring Geometric Feasibility via Affordance-Guided and Self-Consistent MLLMs for Task Planning in Instruction-Following Manipulation
- Department of Computer Science, National Yang Ming Chiao Tung University(计算机科学系,阳明交通大学)
- Department of Computer Science and Information Engineering, National Taiwan University(计算机科学与信息工程系,台湾大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对指令跟随操作中的任务规划,提出QuARC基准并揭示跨模态干扰与几何不可行性问题,通过自洽思维链与可供性引导方法,实现76.7%成功率,显著超越基线。
AI中文摘要:
我们研究了在指令跟随操作中,利用具有上下文学习能力的多模态大语言模型(MLLMs)进行闭环任务规划。我们确定了成功任务规划的四个基本要求:数量估计、可达性分析、相对定位和碰撞避免。然而,现有基准无法全面评估所有这些方面。为解决这一差距,我们引入了QuARC(数量、分析、相对定位、碰撞),这是一个基于食物准备场景的新基准,整合了所有四个挑战。利用QuARC,我们揭示了当前MLLMs的两个主要局限:跨模态干扰和几何不可行性。为解决这些问题,我们采用带自洽性的思维链来缓解跨模态干扰导致的推理损失,并引入可供性预测器来基于几何可行性指导规划。我们的综合评估分析了多个基线的性能,并解释了改进的来源。我们的方法在基准上达到了76.7%的成功率,显著优于ViLa基线(36.7%),且无需额外微调。代码和数据集可在https://hcis-lab.github.io/Affordance-Guided-Self-Consistent-MLLM获取。
英文摘要:
We investigate the use of Multimodal Large Language Models (MLLMs) with in-context learning for closed-loop task planning in instruction-following manipulation. We identify four essential requirements for successful task planning: quantity estimation, reachability analysis, relative positioning, and collision avoidance. However, existing benchmarks fail to support holistic evaluation across all these aspects. To address this gap, we introduce \textbf{QuARC} (Quantity, Analysis, Relative positioning, Collision), a new benchmark based on a food preparation scenario that integrates all four challenges. Using QuARC, we reveal two major limitations of current MLLMs: cross-modal distraction and geometric infeasibility. To tackle these, we adapt Chain-of-Thought with Self-Consistency to mitigate reasoning loss from cross-modal distractions and incorporate an affordance predictor to guide planning based on geometric feasibility. Our comprehensive evaluation analyzes performance across multiple baselines and explains sources of improvement. Our method achieves a 76.7\% success rate on the benchmark, significantly outperforming the ViLa baseline (36.7\%), without requiring additional finetuning. Code and dataset are available at https://hcis-lab.github.io/Affordance-Guided-Self-Consistent-MLLM.