arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

STEP:基于多模态大语言模型的状态感知任务估计与规划,用于人机协作

STEP: State-Aware Task Estimation and Planning with Multi-Modal LLMs for Human-Robot Collaboration

Maitrey Gramopadhye, Prakash Baskaran, Xiao Liu, Songpo Li, Soshi Iba

arXiv 2608.27225首次发表:更新:

发表机构

University of North Carolina at Chapel Hill; Honda Research Institute(北卡罗来纳大学教堂山分校; 本田研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对多模态大语言模型用于人机协作任务规划时的状态理解缺失问题,提出STEP方法,在机器人装配模拟任务中使动作可执行性提升32.8%、最终状态误差降低14.8%。

AI 中文摘要

工业场景中有效的人机协作要求机器人理解人类意图并辅助任务规划,以减少人类工作量。近期研究探索在这类数据稀缺场景中使用多模态大语言模型(MM-LLMs)进行任务规划,利用上下文学习解释用户动作并生成自然语言形式的长 horizon 动作计划。然而,MM-LLMs 固有地缺乏对系统状态的理解,也不跟踪状态转移,常导致生成与预期目标偏离的幻觉动作。此外,用自然语言生成动作计划往往将计划限制在较高层级,导致动作执行时出现歧义。为解决这些局限,我们提出状态感知任务估计器与规划器(STEP),它提示 MM-LLM 显式估计系统状态并预测已执行动作带来的状态转移。通过在生成动作的同时预测未来状态,STEP 确保规划向任务目标收敛,还提供执行预测动作所需的额外辅助参数。我们在模拟环境中使用机器人装配任务评估 STEP,该方法在动作可执行性上比当前最优方法高出 32.8%,在最终状态误差上降低 14.8%。

英文摘要

Effective human-robot collaboration in industrial settings requires robots to understand human intentions and assist with task planning, reducing workload. Recent works have explored the use of Multi-modal Large Language Models (MM-LLMs) for task planning in such data-scarce scenarios, leveraging in-context learning to interpret user actions and generate long-horizon action plans in natural language. However, MM-LLMs inherently lack an understanding of system states and do not track state transitions, often leading to hallucinated actions that deviate from the intended goal. Additionally, generating action plans in natural language tends to limit the generated plans to a high level, introducing ambiguity in action execution. To address these limitations, we propose the State-aware Task Estimator and Planner (STEP), which prompts a MM-LLM to explicitly estimate the state of the system and predict the state transitions resulting from executed actions. By forecasting future states alongside actions, STEP ensures task-convergent planning while also providing additional assistance parameters necessary for executing the predicted actions. We evaluate STEP in a simulated environment using a robot assembly task. Our approach outperforms the state-of-the-art by 32.8% in action executability and 14.8% in final-state error.

CommentsPublished in IEEE International Conference on Robot and Human Interactive Communication (RO-MAN), 2026, 8 pages, 4 figuers, 5 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑