arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向统一的模态无关多模态认知负荷评估框架

Towards a Unified Modality-Agnostic Multimodal Framework for Cognitive Workload Assessment

Stefanos Gkikas, Christian Arzate Cruz, Calvin Joseph, Giorgos Giannakakis, Raul Fernandez Rojas

arXiv 2609.20199首次发表:更新:

AI 中文总结

该研究提出一个统一的模态无关层次Transformer框架,用于认知负荷评估,通过初步实验发现脑电图是最强单模态,且五模态组合在IQ任务上平均分数最高,同时模型大小减少约50%。

AI 中文摘要

认知负荷反映了个体在执行任务时所需的脑力投入,是自适应人机系统设计的核心。利用生物信号测量认知负荷已被广泛研究和记录;然而,关于结合异构生物信号模态用于此目的的研究仍然有限。为了深入探究这一领域,我们开发了一种统一的、模态无关的、基于层次化Transformer的架构,以在单一模型中处理异构生物信号模态。我们在一项初步研究中使用了该框架,评估了五种模态(心电图、皮电活动、呼吸、外周血氧饱和度和脑电图)的所有31种可能组合,并在三种认知不同任务(抽象推理、算术问题解决和游戏任务)上进行了留一受试者交叉验证。在此初步设置中,结果表明:(i)脑电图是最强的单一模态,在IQ、GAME以及合并所有三种任务样本的ALL设置中排名最高;(ii)增加更多模态并不总能持续提升性能;(iii)五种模态的完整组合在IQ上取得了最高的平均分数73.02%,而当将平均分数在四个评估设置(IQ、MATH、GAME和ALL)上取平均时,其平均分数为68.08%;(iv)与晚期融合替代方案相比,所提出的方法将模型大小减少了约50%,同时保持了更低的推理时间。

英文摘要

Cognitive workload reflects the mental effort required during task performance and is central to the design of adaptive human-machine systems. The use of biosignals to measure cognitive workload has been extensively researched and documented; however, studies examining the effects of combining heterogeneous biosignal modalities for this purpose remain limited. To provide insight into this area, we developed a unified, modality-agnostic, hierarchical Transformer-based architecture to process heterogeneous biosignal modalities within a single model. We use this framework in a pilot study evaluating all $31$ possible combinations of five modalities: Electrocardiogram (ECG), Electrodermal Activity (EDA), Respiration (RESP), Peripheral Oxygen Saturation (SpO$_2$), and Electroencephalogram (EEG), under leave-one-subject-out validation across three cognitively distinct tasks: abstract reasoning (IQ), arithmetic problem solving (MATH), and a game task (GAME). In this pilot setting, the results suggest that: (i) EEG is the strongest single modality, ranking highest in IQ, GAME, and the pooled ALL setting, where samples from all three tasks are combined; (ii) adding more modalities does not consistently improve performance; (iii) the full five-modality combination achieves the highest \textit{Average} score of $73.02%$ on IQ and $68.08%$ when the \textit{Average} scores are averaged over the four evaluation settings: IQ, MATH, GAME, and ALL; and (iv) the proposed method reduces model size by approximately $50%$ compared with late-fusion alternatives while maintaining a lower inference time.

CommentsAccepted at the 14th International Conference on Affective Computing and Intelligent Interaction (ACII 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑