arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31898eess.AScs.HCcs.LGcs.SDq-bio.NC

MAESTRO:多模态听觉注意自我中心语音追踪开放语料库

MAESTRO: a Multimodal Auditory-attention Egocentric Speech-TRacking Open corpus

K M Naimul Hassan, Ali Alavi, Donald S. Williamson

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出MAESTRO,首个同时记录EEG、眼动、瞳孔、自我中心视频和头部IMU的多模态听觉注意解码数据集,通过四说话者基准证明融合行为与生理信号优于纯EEG方法。

中文摘要 AI 辅助

人类在嘈杂环境中依赖注视、头部运动和视觉线索来关注说话者,然而听觉注意解码(AAD)主要使用脑电图(EEG)进行研究。我们引入了多模态听觉注意自我中心语音追踪开放(MAESTRO)语料库,这是首个同时记录EEG、眼动注视、瞳孔测量、自我中心视频和头部惯性测量单元(IMU)数据的AAD数据集。MAESTRO包含四个竞争说话者和多种信噪比(SNR)条件下的背景噪声,使得在真实聆听场景下进行注意解码成为可能。通过四说话者注意解码基准测试,我们表明结合行为和生理信号相比仅使用EEG的方法能提高解码性能,为多模态听觉注意解码的未来进展铺平道路。这些发现为多模态AAD的新应用、分析和方法论进步打开了大门。完整数据集可在该https URL公开获取。官方代码库可在该https URL获取。

英文摘要

Humans rely on gaze, head movements, and visual cues to attend to speakers in noisy environments, yet auditory attention decoding (AAD) has been studied primarily using electroencephalography (EEG). We introduce the Multimodal Auditory-attention Egocentric Speech-TRacking Open (MAESTRO) corpus, the first AAD dataset to simultaneously record EEG, eye gaze, pupillometry, egocentric video, and head inertial measurement unit (IMU) data. MAESTRO includes four competing speakers and background noise across multiple signal-to-noise ratio (SNR) conditions, enabling attention decoding under realistic listening scenarios. Through a four-speaker attention decoding benchmark, we show that combining behavioral and physiological signals improves decoding performance over EEG-only approaches, enabling future advances in multimodal auditory attention decoding. These findings open the door to new applications, analyses, and methodological advances in multimodal AAD. The complete dataset is publicly available at https://huggingface.co/datasets/aspire-osu/maestro-eeg-dataset . The official code repository is available at https://github.com/ASPIRE-OSU/MAESTRO .

发表机构

  • The Ohio State University(俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑