基于LoRA的级联多模态融合在医学训练环境中的动作识别
LoRA-Based Cascaded Multimodal Fusion for Action Recognition in Medical Training Environments
浏览论文内容
中文总结 AI 辅助
研究面向医疗训练环境的动作识别,提出基于LoRA的级联多模态融合框架,结合特定模态适应与顺序融合,无需重新训练组件,在两个数据集上评估,结果显示该策略优于单模态模型且性能有竞争力。
中文摘要 AI 辅助
本文提出了一种基于级联低秩适应(LoRA)的多模态融合框架,用于面向医疗保健的训练环境中的动作和活动识别。该架构将参数高效的特定模态适应与顺序融合相结合,使模态能够分阶段集成,而无需重新训练先前学习的组件。该框架不是假设固定的融合结构,而是首先集成关系更紧密的模态,然后纳入其他异构模态,支持跨不同模态数据集的可扩展适应。我们在两个面向医疗保健的训练环境数据集NurViD和护士训练数据集上评估了该框架。初步结果表明,所提出的级联融合策略优于单个模态模型,并相对于先前报告的特定数据集基线提供了有竞争力的性能。总体而言,这些发现表明,基于级联LoRA的融合是一种很有前途的参数高效方法,可用于在医学训练动作和活动识别任务中集成异构模态。
英文摘要
This paper presents a cascaded Low-Rank Adaptation (LoRA)-based multimodal fusion framework for action and activity recognition in healthcare-oriented training environments. The proposed architecture combines parameter-efficient modality-specific adaptation with sequential fusion, enabling modalities to be integrated in stages without retraining previously learned components. Rather than assuming a fixed fusion structure, the framework first integrates more closely related modalities and then incorporates additional heterogeneous modalities, supporting scalable adaptation across datasets with different modality sets.We evaluate the framework on two healthcare-oriented training environment datasets: NurViD and the Nurse Training dataset. Across these datasets, preliminary results suggest that the proposed cascaded fusion strategy improves over individual modality models and provides competitive performance relative to previously reported dataset-specific baselines. Overall, these findings indicate that cascaded LoRA-based fusion is a promising parameter-efficient approach for integrating heterogeneous modalities in medical training action and activity recognition tasks. github: https://github.com/anonymous0-ai/LoRA-Based-Cascaded-Multimodal-Fusion-.git.
发表机构
- Quince(Quince公司)
- Vanderbilt University(范德堡大学)
机构由 AI 辅助整理,请以论文原文为准。