arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于听觉注意力解码的端到端马尔可夫状态序列学习

End-to-End Markov State Sequence Learning for Auditory Attention Decoding

Yushan Yashengjiang, Jie Zhang, Miao Sun, Huadong Liang, Xin Li, Zhen-hua Ling

arXiv 2607.18614首次发表:更新:

发表机构

NERC-SLIP, University of Science and Technology of China; School of Information and Communication Engineering, Guangzhou Maritime University; Artificial Intelligence Research Institute, iFLYTEK Company, Ltd.; School of Information Science and Technology, University of Science and Technology of China(中国科学技术大学NERC-SLIP; 广州航海学院信息与通信工程学院; 科大讯飞股份有限公司人工智能研究院; 中国科学技术大学信息科学技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对听觉注意力解码中多数模型为独立短窗口分类器的问题,提出基于条件随机场的端到端马尔可夫AAD框架,引入ESCNet,通过联合优化目标学习注意力状态序列,实验证明该方法在多个数据集上优于传统孤立窗口分类。

AI 中文摘要

听觉注意力解码(AAD)从脑电图(EEG)等神经反应中识别听众关注的说话者,是神经导向助听器中的关键算法。但多数神经AAD模型是独立短窗口分类器,忽略了听觉注意力的时间持续性。本文提出基于条件随机场(CRF)的端到端马尔可夫AAD框架,在两状态注意力先验下训练窗口级神经发射。该框架将AAD主干的对数几率视为马尔可夫发射,从标准隐马尔可夫模型初始化学习转移率,联合优化交叉熵和CRF目标。还引入了ESCNet作为EEG-语音相关主干。实验表明,在动态AVGC数据集上,CRF训练优于事后隐马尔可夫模型平滑;在静态KUL和USTC数据集上,相比固定速率事后隐马尔可夫模型基线,因果解码分别提高了5.6%和2.0%,显示出将AAD学习为注意力状态序列优于孤立窗口分类。

英文摘要

Auditory attention decoding (AAD) identifies the speaker a listener attends to from neural responses like electroencephalography (EEG), making it a key algorithm in neuro-steered hearing aids. However, most neural AAD models are trained as independent short-window classifiers, despite auditory attention being a temporally persistent cognitive state and short-window EEG--audio evidence often being noisy and ambiguous. We propose an end-to-end Markov AAD framework based on conditional random field (CRF) that trains window-level neural emissions under a two-state attention prior. The framework treats the logits of any AAD backbone as Markov emissions, learns the transition rate from a standard HMM initialization, and jointly optimizes cross-entropy and CRF objectives, allowing temporal continuity to guide representation learning rather than merely smoothing predictions after training. We also introduce ESCNet, an EEG--speech correlation backbone that preserves time-aligned features and converts the difference between two mean Pearson correlations into state logits. We evaluate the framework with four emission backbones spanning correlation-based, convolutional, recurrent, and attention-based designs. On the dynamic AVGC dataset, CRF training generally outperforms post-hoc HMM smoothing; with ESCNet, it achieves $86.5\%$ causal and $92.4\%$ non-causal accuracy using $1$s windows. On the static KUL and USTC datasets, it improves causal decoding over fixed-rate post-hoc HMM baselines by $5.6\%$ and $2.0\%$, respectively, showing the superiority of learning AAD as attention state sequence over isolated-window classification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑