BAT-CLIP:脑、音频与文本的三模态对齐
BAT-CLIP: Trimodal Alignment of Brain, Audio and Text
浏览论文内容
中文总结 AI 辅助
提出BAT-CLIP,首个针对iEEG的CLIP式三模态对齐框架,联合对齐神经嵌入与音频和文本锚点,在Podcast基准上比双模态基线更稳健。
中文摘要 AI 辅助
从大脑中解码和解读自然语音越来越依赖于与预训练的语音和语言表示空间的对齐。然而,当前的CLIP式脑-语音对齐将神经活动锚定到单一模态——音频或文本——尽管大脑本质上进行多模态语音处理。这导致了一种权衡:音频锚定保留了时间结构但削弱了语言可分离性,而文本锚定捕获了语义却丢弃了声学细节。我们提出了BAT-CLIP,这是首个用于iEEG的CLIP式三模态对齐框架,它将神经嵌入与预训练的音频和文本锚点在共享的冻结音频-文本流形中联合对齐。在自然语音Podcast基准上,BAT-CLIP比双模态CLIP基线产生了更稳健的表示。我们还强调了使用自监督基础模型进行CLIP训练的重要性。
英文摘要
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech processing. This induces a trade-off: audio anchoring preserves temporal structure but weakens linguistic separability, while text anchoring captures semantics yet discards acoustic detail. We propose BAT-CLIP, the first CLIP-style trimodal alignment framework for iEEG that jointly aligns neural embeddings to both pretrained audio and text anchors in a shared, frozen audio-text manifold. On the naturalistic Podcast benchmark, BAT-CLIP yields more robust representations than bimodal CLIP baselines. We also highlight the importance of using self-supervised foundation models for CLIP training.
发表机构
- Yonsei University(延世大学)
- Seoul National University(首尔国立大学)
- Dartmouth College(达特茅斯学院)
机构由 AI 辅助整理,请以论文原文为准。