arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在人机协作中学习规划:用于自适应交互的多模态强化学习

Learning to Plan in Human-Robot Collaboration: Multimodal Reinforcement Learning for Adaptive Interaction

Afagh Mehri Shervedani, Siyu Li, Natawut Monaikul, Bahareh Abbasi, Barbara Di Eugenio, Miloš Žefran

arXiv 2609.25274首次发表:更新:

发表机构

University of Illinois Chicago; California State University Channel Islands(伊利诺伊大学芝加哥分校; 加州州立大学海峡群岛分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种强化学习方法,自动生成机器人在家庭环境中协助用户定位物体的多模态交互策略,实验表明该方法具有高可用性和可扩展性。

AI 中文摘要

面向老年人和残疾人的机器人助手需要与用户有效地执行协作任务。这些系统的核心组件是一个交互管理器,其职责是观察和评估任务,推断人类的状态及其意图,以便机器人选择最佳行动方案。由于该领域数据的稀疏性,此类多模态系统的策略通常由人工设计;随着交互复杂性的增长,这一过程变得不可扩展。本文提出了一种强化学习(RL)方法,用于自动生成机器人的多模态策略。我们的系统聚焦于一个现实场景:机器人在家庭环境中协助用户定位物体,管理包括语言和物理动作在内的多模态信号,以选择最佳行动。与传统对话系统不同,我们的智能体使用基于人类数据的模拟器进行训练,并且能够处理多种模态。我们使用一个简单的高层奖励函数,无需微调,并强制执行一些前提条件以加速训练过程。一项在真实环境中评估该系统的人类研究展示了有前景的结果,表明其具有高可用性和有效的任务完成度。这种基于RL的方法为设计多模态人机协作中的交互管理器提供了一种可扩展且可解释的替代方案。

英文摘要

Robot assistants for older adults and people with disabilities need to perform collaborative tasks with users effectively. The core component of these systems is an interaction manager whose job is to observe and assess the task and infer the state of the human and their intent for the robot to choose the best course of action. Due to the sparseness of the data in this domain, the policy for such multimodal systems is often crafted by hand; as the complexity of interactions grows, this process is not scalable. This paper proposes a reinforcement learning (RL) approach to automatically generate the multimodal policy of the robot. Our system focuses on a realistic scenario where a robot assists a user in locating objects within a home environment, managing multimodal signals, including language and physical actions, to select the best action. In contrast to traditional dialog systems, our agent is trained with a simulator that uses human data and can deal with multiple modalities. We use a simple high-level reward function that needs no fine-tuning and enforce some preconditions to speed up the training process. A human study evaluating the system in a real-world setting demonstrates promising results, indicating high usability and effective task completion. This RL-based approach offers a scalable and interpretable alternative for designing interaction managers in multimodal human-robot collaborations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑