AI 中文总结
提出置信度感知多模态融合框架,融合四种模态并利用动态融合与平衡学习,在UR3平台上实现91.86%的意图识别准确率,支持稳定人机协作。
AI 中文摘要
提出了一种置信度感知的多模态融合框架(CAMF),以实现工业人机协作中可靠的人类意图预测。该框架融合了四种异构模态,包括物体6D位姿、注视、骨骼运动和基于IMU的手部运动。它将置信度趋势驱动的动态融合机制嵌入到BiLSTM中,根据实时模态可靠性自适应地平衡双向时间特征。进一步采用置信度引导的平衡学习策略与置信度冻结机制相结合,动态调整网络梯度,抑制低质量模态的噪声并减轻跨模态学习偏差。搭建了基于UR3协作机器人的物理平台进行实验验证。对比结果表明,所提方法的意图识别准确率达到91.86%,在整体性能和稳定性上优于现有的多模态融合方法。在低光照和部分遮挡干扰下,该方法仍能保持令人满意的准确性。在实际装配任务中,该框架能够实现主动且稳定的人机协作,具有较强的环境适应性。
英文摘要
A confidence-aware multimodal fusion framework (CAMF) is proposed to realize reliable human intention prediction for industrial human-robot collaboration. This framework fuses four heterogeneous modalities including object 6D pose, gaze, skeletal motion and IMU-based hand motion. It embeds a confidence-trend-driven dynamic fusion mechanism into BiLSTM to adaptively balance bidirectional temporal features according to real-time modality reliability. A confidence-guided balanced learning strategy combined with a confidence freezing mechanism is further adopted to adjust network gradients dynamically, suppress noise from low-quality modalities and mitigate cross-modal learning bias. A physical platform based on the UR3 collaborative robot is built for experimental validation. Comparative results show that the proposed method reaches an intention recognition accuracy of 91.86% and outperforms existing multimodal fusion approaches in overall performance and stability. It also maintains satisfactory accuracy under low light and partial occlusion interference. In practical assembly tasks, the framework enables proactive and stable human-robot cooperation with strong environmental adaptability.