arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于机器人的智能云边缘多模态交互系统

An Intelligent-Cloud Edge Multimodal Interaction System for Robots

Zihan Guo, Xiaoqi Li

arXiv 2607.14675首次发表:更新:

发表机构

Hainan University(海南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对复杂环境下资源受限机器人的交互问题,提出云边缘多模态交互框架,集成增强YOLO手势检测器与LLM、VLM智能体,改进手势检测方法,经实验验证该系统在手势检测精度、任务成功率及用户满意度方面表现良好,证明了方法的可行性。

AI 中文摘要

在复杂环境中实现强大的人机交互,需要在有限的机载计算资源下进行精确的手势感知、语义场景理解和可靠的任务规划。本文提出一种云边缘多模态交互框架,集成了基于增强YOLO的手势检测器以及协同的大语言模型(LLM)和视觉语言模型(VLM)智能体。该检测器在颈部引入卷积块注意力模块(CBAM),并用距离交并比(DIoU)损失取代基线边界框回归目标,改进了复杂背景下小或部分遮挡手势的特征辨别和定位。云层执行手势检测、场景理解、多模态融合和动作规划,TonyPi机器人本地处理数据采集、通信、动作执行和反馈。在公共手势数据集和自定义数据集上的实验表明,YOLO-DC的精度值分别为98.9%和95.0%,mAP@0.5值分别为90.7%和92.7%。系统级评估得出单动作、复合动作和视觉相关任务的成功率分别为95%、88%和82%。30名参与者的评估得出总体平均满意度评分为3.69(满分5分)。这些结果证明了将精细的手势检测与多模态智能体相结合用于资源受限机器人交互的可行性。

英文摘要

Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources. This paper presents a cloud-edge multimodal interaction framework that integrates an enhanced YOLO-based gesture detector with coordinated large language model (LLM) and vision-language model (VLM) agents. The proposed detector, incorporates the Convolutional Block Attention Module (CBAM) into the neck and replaces the baseline bounding-box regression objective with Distance-IoU (DIoU) loss. These modifications improve feature discrimination and localization for small or partially occluded gestures in complex backgrounds. The cloud layer performs gesture detection, scene understanding, multimodal fusion, and action planning, whereas the TonyPi robot locally handles data acquisition, communication, action execution, and feedback. Experiments on a public gesture dataset and a custom dataset show that YOLO-DC achieves precision values of 98.9% and 95.0%, with mAP@0.5 values of 90.7% and 92.7%, respectively. System-level evaluation yields success rates of 95%, 88%, and 82% for single-action, composite-action, and vision-dependent tasks. A 30 participant evaluation yields an overall mean satisfaction score of 3.69 out of 5. These results demonstrate the feasibility of combining refined gesture detection with multimodal agents for resource-constrained robotic interaction.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑