ProAct-VLM:具有连续感知反馈的故障前视觉语言任务重规划
ProAct-VLM: Pre-Failure Vision-Language Task Replanning with Continuous Perception Feedback
浏览论文内容
中文总结 AI 辅助
ProAct-VLM提出一种自适应任务规划框架,通过实时感知反馈循环集成视觉语言模型,在故障发生前持续监控环境并重规划,提高动态长时程操作任务的成功率和效率。
中文摘要 AI 辅助
长时程机器人任务容易受到意外环境变化的影响,这些变化可能使规划的动作失效或不安全。为了解决这个问题,机器人必须在变化发生时检测到它们,理解其影响,并相应地调整动作。传统的基于规则的决策流程在开放世界条件下是脆弱的,因为它们针对特定场景手工调整,缺乏泛化能力。视觉语言模型(VLMs)提供了一种有前景的替代方案,因为它们结合了广泛的世界知识与统一的视觉-文本推理,使它们能够跨多样场景泛化并生成准确、有根据的任务计划。然而,为了在动态真实世界环境中有效部署,VLMs必须嵌入能够处理不确定性和环境变化的框架中。现有框架广泛地以反应式方式处理这个问题,仅在执行失败或任务后检查后触发重规划,存在动作失败的风险。一些方法在动作前验证条件,但这些离散检查会错过执行期间发生的变化。为了解决这个问题,我们提出了ProAct-VLM,一个自适应的、物理基础的任务规划框架,它将VLMs集成在实时感知-反馈循环中。ProAct-VLM持续监控环境,并在检测到相关变化时立即重新规划,从而在故障发生前实现适应。针对多个基线和不同VLM骨干的评估表明,我们的框架在动态、长时程操作任务中提高了成功率和效率。项目页面:此https URL。
英文摘要
Long-horizon robotic tasks are vulnerable to unexpected environmental changes that can render planned actions ineffective or unsafe. To address this, robots must detect such changes as they occur, interpret their impact, and adjust their actions accordingly. Traditional rule-based decision-making pipelines are brittle in open-world conditions, as they are hand-tuned for specific scenarios and lack generalization. Vision-Language Models (VLMs) offer a promising alternative as they combine broad world knowledge with unified visual--text reasoning, enabling them to generalize across diverse scenarios and generate accurate, grounded task plans. However, for effective deployment in dynamic real-world settings, VLMs must be embedded into frameworks capable of handling uncertainty and environmental changes. Existing frameworks broadly address this reactively, triggering replanning only after execution failures or post-task checks, risking failed actions. Some methods verify conditions before actions, but these discrete checks miss changes occurring during execution. To address this, we present ProAct-VLM, an adaptive, physically grounded task planning framework that integrates VLMs within a real-time perception--feedback loop. ProAct-VLM continuously monitors the environment and re-plans as soon as relevant changes are detected, enabling adaptation before failure occurs. Evaluations against multiple baselines and across different VLM backbones show that our framework improves both success rates and efficiency in dynamic, long-horizon manipulation tasks. Project page: https://github.com/moured/ProAct-VLM
发表机构
- Khalifa University(哈利法大学)
- Sereact GmbH(Sereact公司)
机构由 AI 辅助整理,请以论文原文为准。