基于多智能体强化学习的联邦多模态人体活动识别
FedMHAR: Federated Multimodal Human Activity Recognition using Multi-Agent Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
针对多模态HAR中传感器权重固定和联邦学习优化问题,提出基于MARL的集中式框架及联邦扩展FedMHAR,通过PPO智能体动态分配融合权重和BiFL-PPO双向优化,在MEx和UTD数据集上取得更高准确率并降低成本。
中文摘要 AI 辅助
从异构可穿戴传感器进行人体活动识别(HAR)是健康物联网(IoHT)的基础,支持康复、老年护理和智能医疗。现有的多模态融合方法通常为传感器流分配固定的相等权重,忽视了模态重要性、采集成本和传感器质量的差异,这些差异可能因运动、佩戴位置不当或临时遮挡而变化。我们提出了一种基于多智能体强化学习的自适应且成本感知的多模态HAR框架,用于集中式HAR,并将其扩展至联邦学习,称为FedMHAR。在集中式设置中,多模态融合被建模为协作式多智能体强化学习(MARL)问题,每个传感模态分配一个基于PPO的智能体,学习每个样本的融合权重,使模型能够强调信息丰富的模态,同时在更便宜的替代方案提供足够信息时降低昂贵传感器的权重。在联邦设置中,我们引入了BiFL-PPO,一种双向联邦优化策略,其中服务器端PPO策略学习客户端特定的信任权重,并将其反馈以调整本地学习率和近端正则化。与轮级优化不同,BiFL-PPO使用密集的批次级奖励,以在异构客户端数据下提供更频繁的反馈和稳定的训练。在MEx康复和UTD多模态人体动作数据集上的评估表明,集中式框架分别达到87.30%和94.98%的准确率,优于传统融合方法和最先进的HAR模型。FedMHAR在联邦设置中达到79.74%和77.49%,持续超越FedAvg、FedProx、FedBN、FedNova和AdaFedProx,同时提供更稳定的性能并降低传感器采集成本。
英文摘要
Human Activity Recognition (HAR) from heterogeneous wearable sensors is fundamental to the Internet of Health Things (IoHT), supporting rehabilitation, elderly care, and smart healthcare. Existing multimodal fusion methods often assign fixed equal weights to sensor streams, overlooking differences in modality importance, acquisition cost, and sensor quality, which can vary due to movement, incorrect placement, or temporary blockage. We propose an adaptive and cost-aware multimodal HAR framework based on multi-agent reinforcement learning for centralized HAR and extend it to federated learning as FedMHAR. In the centralized setting, multimodal fusion is formulated as a cooperative Multi-Agent Reinforcement Learning (MARL) problem, where each sensing modality is assigned a PPO-based agent that learns per-sample fusion weights, enabling the model to emphasize informative modalities while down-weighting costly sensors when cheaper alternatives provide sufficient information. In the federated setting, we introduce BiFL-PPO, a bidirectional federated optimization strategy in which a server-side PPO policy learns client-specific trust weights and feeds them back to adapt local learning rates and proximal regularization. Unlike round-level optimization, BiFL-PPO uses dense batch-level rewards for more frequent feedback and stable training under heterogeneous client data. Evaluation on the MEx Rehabilitation and UTD Multimodal Human Action datasets shows that the centralized framework achieves 87.30% and 94.98% accuracy, respectively, outperforming conventional fusion methods and state-of-the-art HAR models. FedMHAR achieves 79.74% and 77.49% in the federated setting, consistently surpassing FedAvg, FedProx, FedBN, FedNova, and AdaFedProx, while providing more stable performance and reducing sensor acquisition cost.
发表机构
- Indian Statistical Institute Kolkata(加尔各答印度统计研究所)
- Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。