发表机构
University of Notre Dame; Washington State University(圣母大学; 华盛顿州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大型推理语言模型用于机器人任务规划时的推理可靠性与计算浪费问题,提出SafeInferCom框架,通过验证器引导的生成中干预提升规划成功率、加速错误修正并减少token使用,经多模型、多领域及真实机械臂验证有效。
AI 中文摘要
大型推理语言模型(LRLM)可实现机器人任务规划的多步推理,但持续推理可能覆盖有效中间计划或未解决约束违反问题,降低规划可靠性并浪费推理时计算资源。我们开发了一种推理时监控器,可在不干扰原始解码轨迹的情况下暴露并验证中间计划。基于该监控器,我们提出SafeInferCom,这是一种形式化验证器引导的框架,可保留有效中间计划并在生成过程中指导错误修正。在多个LRLM和规划领域的实验显示,单次推理下存在推理-响应不一致和自我修正能力有限的问题。SafeInferCom相比单次推理提高了规划成功率并加速了错误修正;与迭代优化结合时,相比单独优化进一步提高了成功率,同时减少了token使用量。我们还在VirtualHome中评估了SafeInferCom,并提供了真实世界机械臂演示。
英文摘要
Large Reasoning Language Models (LRLMs) enable multi-step reasoning for robotic task planning, but continued reasoning can overwrite valid intermediate plans or leave constraint violations unresolved, reducing planning reliability and wasting inference-time computation. We develop an inference-time monitor that exposes and verifies intermediate plans without disrupting the original decoding trajectory. Building on this monitor, we propose SafeInferCom, a formal verifier-guided framework that preserves valid intermediate plans and directs error correction during generation. Experiments across multiple LRLMs and planning domains reveal reasoning-response inconsistency and limited self-correction under one-shot inference. SafeInferCom improves planning success and accelerates error correction relative to one-shot inference. When combined with iterative refinement, it further improves success while reducing token usage compared with refinement alone. We additionally evaluate SafeInferCom in VirtualHome and provide a real-world robotic-arm demonstration.
CommentsVideo: https://youtu.be/dbU7WskCNgY