面向养老机构服务机器人的多模态语言模型驱动的交互与陪伴
Multimodal-Language-Model-Driven Interaction and Companionship for Service Robots in Elderly-Care Facilities
浏览论文内容
中文总结 AI 辅助
该研究针对养老机构服务机器人缺乏综合陪伴与安全能力的问题,提出整合主动视觉跟人、LLM语音交互及VLM安全监测的智能陪伴机器人系统,实验验证其跟随交互与跌倒检测的有效性。
中文摘要 AI 辅助
服务机器人正越来越多地部署在养老机构中,以减轻护理人员的工作量并提升日常护理的质量。然而,现有多数研究聚焦于孤立的服务功能,缺乏持续陪伴、自然交互及安全监测的综合能力。本文提出一种智能陪伴机器人系统,该系统整合了主动视觉人跟、基于实时大语言模型(LLM)的用于意图理解与任务执行的语音交互,以及基于视觉语言模型(VLM)的用于跌倒检测与异常姿态评估的安全监测。感知层确保稳健的人体跟踪,并在遮挡或突发运动时使用主动云台将用户保持在视野内。交互层中,大语言模型解释口头请求并将其映射为机器人动作,实现护送与语义导航。同时,基于VLM的安全智能体持续分析视觉观测结果,以检测跌倒相关或异常姿态,并在必要时触发应急响应。实验结果表明,该系统具备可靠跟随与交互的能力,同时能有效检测潜在跌倒以保障用户安全。
英文摘要
Service robots are increasingly deployed in elderly-care facilities to alleviate caregiver workload and enhance the quality of daily care. However, most existing studies focus on isolated service functions and lack integrated capabilities for continuous companionship, natural interaction, and safety monitoring. In this paper, we present an intelligent companion robot system that unifies active visual human-following, real-time LLM-driven speech interaction for intent understanding and task execution, and VLM-based safety monitoring for fall detection and abnormal posture assessment. The perception layer ensures robust human tracking and uses an active gimbal to maintain the user in view during occlusions or abrupt movements. At the interaction layer, a Large Language Model interprets spoken requests and maps them to robot actions, enabling escorting and semantic navigation. Simultaneously, a VLM-based safety agent continuously analyzes visual observations to detect fall-related or abnormal postures and triggers emergency responses when necessary. Experimental results demonstrate the system's ability to reliably follow and interact with humans, while effectively detecting potential falls to ensure user safety.
发表机构
- National Cheng Kung University (NCKU)(成功大学)
机构由 AI 辅助整理,请以论文原文为准。