AI 中文总结
研究针对助听器中听觉注意力解码因数据有限面临的挑战,利用扩散概率模型生成合成语音诱发EEG数据进行数据增强,实验证明该方法能显著提高AAD性能,凸显其减轻训练数据限制及增强模型鲁棒性的潜力。
AI 中文摘要
有限的训练数据限制了助听器中用于听觉注意力解码(AAD)的深度学习模型。AAD利用脑电图(EEG)数据解码听众的注意力,以实时跟踪特定声源。然而,由于现实世界中语音诱发EEG数据稀缺,在助听器典型的短时间窗口(<=1秒)内实现高AAD性能具有挑战性。为解决此问题,研究了扩散概率模型(DPMs)来生成合成语音诱发EEG数据。DPMs通过去噪过程学习潜在复杂数据结构,能生成适合数据增强的逼真样本。评估了合成EEG数据在注意力位置(LoA)分类任务中增强数据集的应用。实验表明DPMs能生成逼真EEG信号,与仅用实测EEG数据训练的模型相比,纳入合成数据显著提高了AAD性能(p<0.05)。这些结果凸显了基于扩散的数据增强减轻训练数据限制及提高短窗口AAD模型在助听器应用中鲁棒性的潜力。
英文摘要
Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs). AAD uses electroencephalogram (EEG) data to decode listener's attention, enabling real-time tracking of specific sound sources. However, achieving high AAD performance with short time windows typical in HAs (<=1s) is challenging due to the scarcity of real-world speech-evoked EEG data. To address this issue, we investigate diffusion probabilistic models (DPMs) for generating synthetic speech-evoked EEG data. DPMs learn the underlying complex data structure through a denoising process and can generate realistic samples suitable for data augmentation. We evaluate the use of synthetic EEG data for augmenting datasets in locus-of-attention (LoA) classification tasks. Our experiments demonstrate that DPMs can generate realistic EEG signals and that incorporating synthetic data significantly improves AAD performance compared to models trained solely on measured EEG data (p<0.05). These results highlight the potential of diffusion-based data augmentation to mitigate training data limitations and improve the robustness of short-window AAD models in HA applications.