EMG-FM-Bench:面向肌电信号的基础模型迁移与适应的综合基准
EMG-FM-Bench: A Comprehensive Benchmark for Foundation Model Transfer and Adaptation on Electromyography
浏览论文内容
中文总结 AI 辅助
本基准整合20个数据集和9个预训练模型,系统评估了基础模型在肌电信号上的迁移与适应,发现预训练收益因模型而异,微调显著影响性能,且新用户适应仅部分恢复性能。
中文摘要 AI 辅助
基础模型(FMs)正越来越多地被开发用于通用时间序列和生理信号,然而它们向下游生理任务的迁移能力仍鲜为人知。这一问题在肌电信号(EMG)领域尤为棘手,因为信号分布在用户、传感配置、采集硬件和下游任务之间存在显著差异。我们提出了EMG-FM-Bench,一个用于研究基础模型在肌电信号上迁移与适应的系统性基准。EMG-FM-Bench整合了20个公共数据集,包含超过100万个肌电信号片段,并围绕四个问题评估了九个预训练基础模型:预训练模型在冻结或完全微调时的表现如何;与从零训练同一模型相比,预训练带来多大帮助;模型在有限标注数据下对新用户的泛化能力如何;以及不同肌电任务之间的性能变化。在整个基准中,线性探测为预训练表示提供了有用信息,但完全微调能显著改变下游肌电性能。将每个预训练模型与从零训练的同一模型进行比较,发现预训练的收益在不同模型间差异显著,并非普遍存在。当模型在新用户上评估时性能下降,而五次样本适应在70.2%的模型-数据集组合中提升了宏F1分数,但仅部分恢复了损失的性能。模型性能在上肢和下肢分类之间高度一致,并与连续的肌电到文本解码保持强相关。综合来看,这些结果系统地揭示了预训练时间序列模型何时能有效迁移到肌电信号,以及其性能如何依赖于微调、用户变异和下游任务。
英文摘要
Foundation models (FMs) are increasingly being developed for general time series and physiological signals, yet their transferability to downstream physiological tasks remains poorly understood. This question is particularly challenging for electromyography (EMG), where signal distributions vary substantially across users, sensing configurations, acquisition hardware, and downstream tasks. We introduce EMG-FM-Bench, a systematic benchmark for studying foundation-model transfer and adaptation on EMG. EMG-FM-Bench unifies 20 public datasets with over 1 million EMG segments and evaluates nine pretrained foundation models across four questions: how pretrained models perform when frozen or fully fine-tuned, how much pretraining helps compared with training the same model from scratch, how well models generalize to new users with limited labeled data, and how performance changes across different EMG tasks. Across the benchmark, linear probing provides useful information about pretrained representations, but full fine-tuning can substantially change downstream EMG performance. Comparing each pretrained model with the same model trained from scratch shows that the benefit of pretraining varies substantially across models and is not universal. Performance decreases when models are evaluated on new users, while five-shot adaptation improves macro-F1 in 70.2% of evaluated model-dataset combinations but recovers only part of the lost performance. Model performance is highly consistent between upper- and lower-limb classification and remains strongly correlated with continuous EMG-to-text decoding. Together, these results provide a systematic view of when pretrained time-series models transfer effectively to EMG and how their performance depends on fine-tuning, user variation, and downstream task.
发表机构
- University of Georgia(佐治亚大学)
- University of Oklahoma(俄克拉荷马大学)
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。