arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向主动式机器人的认知动作推理:基于以人为中心的多模态观测

Cognitive Action Reasoning for Proactive Robots from Human-Centered Multimodal Observations

Zhihao Gu, Kechao Zhu, Yuanfeng Wu, Mohan Liu, Ankit Kumar Shaw, ChenDong Hong, Xuanyu Chen, Dengchen Mei, Xu Tianyi, Lin Wang

arXiv 2609.23486首次发表:更新:

发表机构

Nanyang Technological University (NTU)(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对机器人缺乏主动决策能力的问题,提出ProRobo认知推理框架及ProAction多模态数据集,并构建MMC2Act模型,显著提升机器人从多模态线索推理高层动作的性能。

AI 中文摘要

在以人为本的环境中运行的机器人通常被设计为执行明确的指令,且大多数机器人学习数据集同样将观测与任务指令或低级动作配对。尽管近期工作已开始探索主动式具身辅助,但现有资源针对不同的场景和动作层级,使得现实世界中以人为中心的多模态决策仍未被充分探索。我们将此问题形式化为《主动式机器人动作推理》(ProRobo),这是一个上游认知决策问题,其中机器人必须基于多模态的人类与环境线索决定采取何种动作,而无需明确的动作指令。为支持ProRobo,我们引入了《ProAction》,一个真实世界的多模态数据集,包含10K个样本,涵盖五个常见场景中12个日常生活情境的视觉观测、音频信号和文本输入。为构建认知基础的高层动作监督,我们开发了一个两阶段的人机协同流程,结合了基于评价的候选生成与基于情感心智理论的人类精炼,明确将关于人类状态、紧迫性、可行性和潜在风险的情境判断纳入动作标注中。基于此监督,我们对代表性的多模态大语言模型(MLLMs)进行了基准测试,并引入了《MMC2Act》,一个参考模型,隐式学习从多模态观测到认知基础的高层动作的映射。跨模态设置、主体不相交泛化、跨数据集迁移及人类评估的实验表明,通用MLLMs在从多模态线索主动推理高层动作方面存在困难,而在《ProAction》上进行训练则显著提升了性能。

英文摘要

Robots operating in human-centered environments are typically designed to execute explicit instructions, and most robot-learning datasets likewise pair observations with task instructions or low-level actions. Although recent work has begun to explore proactive embodied assistance, existing resources target different settings and action levels, leaving real-world human-centered multimodal decision-making underexplored. We formulate this problem as \textit{Proactive Robot Action Reasoning} (\textit{ProRobo}), an upstream cognitive decision problem in which a robot must determine which action to take based on multimodal human and environmental cues without explicit action instructions. To support ProRobo, we introduce \textit{ProAction}, a real-world multimodal dataset containing 10K samples of visual observations, audio signals, and text inputs across 12 daily-life scenarios in five common scenes. To construct cognitively grounded high-level action supervision, we develop a two-stage human-in-the-loop pipeline that combines appraisal-guided candidate generation with Affective Theory-of-Mind-guided human refinement, explicitly incorporating contextual judgment about human states, urgency, feasibility, and potential risk into action annotation. Based on this supervision, we benchmark representative Multimodal Large Language Models (MLLMs) and introduce \textit{MMC2Act}, a reference model that implicitly learns the mapping from multimodal observations to cognitively grounded high-level actions. Experiments across modality settings, subject-disjoint generalization, cross-dataset transfer, and human evaluation show that general-purpose MLLMs struggle with proactively reasoning high-level actions from multimodal cues, whereas training on \textit{ProAction} substantially improves performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑