arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用扩散生成模型解决听觉注意力解码中的数据有限问题

Addressing Limited Data in Auditory Attention Decoding with Diffusion Generative Models

David Rannaleet, Victor Gunnarsson, Bo Bernhardsson, Martin A. Skoglund, Emina Alickovic

arXiv 2607.18345首次发表:更新:

AI 中文总结

研究针对助听器中听觉注意力解码因数据有限面临的挑战,利用扩散概率模型生成合成语音诱发EEG数据进行数据增强,实验证明该方法能显著提高AAD性能,凸显其减轻训练数据限制及增强模型鲁棒性的潜力。

AI 中文摘要

有限的训练数据限制了助听器中用于听觉注意力解码(AAD)的深度学习模型。AAD利用脑电图(EEG)数据解码听众的注意力,以实时跟踪特定声源。然而,由于现实世界中语音诱发EEG数据稀缺,在助听器典型的短时间窗口(<=1秒)内实现高AAD性能具有挑战性。为解决此问题,研究了扩散概率模型(DPMs)来生成合成语音诱发EEG数据。DPMs通过去噪过程学习潜在复杂数据结构,能生成适合数据增强的逼真样本。评估了合成EEG数据在注意力位置(LoA)分类任务中增强数据集的应用。实验表明DPMs能生成逼真EEG信号,与仅用实测EEG数据训练的模型相比,纳入合成数据显著提高了AAD性能(p<0.05)。这些结果凸显了基于扩散的数据增强减轻训练数据限制及提高短窗口AAD模型在助听器应用中鲁棒性的潜力。

英文摘要

Limited training data constrains deep learning models for Auditory Attention Decoding (AAD) in hearing aids (HAs). AAD uses electroencephalogram (EEG) data to decode listener's attention, enabling real-time tracking of specific sound sources. However, achieving high AAD performance with short time windows typical in HAs (<=1s) is challenging due to the scarcity of real-world speech-evoked EEG data. To address this issue, we investigate diffusion probabilistic models (DPMs) for generating synthetic speech-evoked EEG data. DPMs learn the underlying complex data structure through a denoising process and can generate realistic samples suitable for data augmentation. We evaluate the use of synthetic EEG data for augmenting datasets in locus-of-attention (LoA) classification tasks. Our experiments demonstrate that DPMs can generate realistic EEG signals and that incorporating synthetic data significantly improves AAD performance compared to models trained solely on measured EEG data (p<0.05). These results highlight the potential of diffusion-based data augmentation to mitigate training data limitations and improve the robustness of short-window AAD models in HA applications.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑