发表机构
Ohio State University(俄亥俄州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
NEUROTOKEN通过条件流匹配联合建模源与方向听觉注意解码,利用共享EEG前端和似然比评分,在多个数据集上显著提升准确率并缩小方差。
AI 中文摘要
在嘈杂房间中识别听者正在关注哪位说话者——即鸡尾酒会问题——是下一代助听器和脑机接口缺失的关键要素:它告诉设备应放大谁的声音。听觉注意解码(AAD)从脑电图(EEG)中读取这一答案,但现有文献分裂为互不关联的部分:方向性AAD分类侧别但不将侧别映射到语音流;基于回归的源AAD通过单一皮尔逊相关系数对候选语音流排序,该系数在真实设备所需的1-5秒窗口内固有噪声较大;而包络重建没有原生AAD规则。我们认为正确的对象不是任何单一统计量,而是给定脑电图下被注意包络的条件似然,并通过NEUROTOKEN将其实用化:一个单一网络,其三个输出头共享一个脑电图前端,其中一个条件流匹配头(ATTUNEFLOW)通过积分速度残差似然比来对候选者评分。两个推理时集成——QUADTRACK(四个互补统计量)和ENV-FLOW(z归一化的QUADTRACK+ATTUNEFLOW)——免费吸收各统计量的失败模式。在KU Leuven、DTU和NJU数据集上,5秒窗口下,ATTUNEFLOW将逐段源AAD相较于最强非生成基线提升9%-16%,并将跨受试者方差缩小约3倍;试验级融合在三个数据集中的两个上超过93%。在并行复现中,我们显示经典的95%-97%方向AAD数字在严格的试验不相交协议下下降17%-45%,从而阐明了真实上限以及为何需要基于似然的公式。
英文摘要
Identifying which speaker a listener is attending to in a noisy room -- the cocktail-party problem -- is the missing ingredient for next-generation hearing aids and brain-computer interfaces: it tells the device whose voice to amplify. Auditory attention decoding (AAD) reads this answer from EEG, but the literature splits into disconnected pieces: directional-AAD classifies side but does not map side to stream; regression-based source-AAD ranks candidate streams by a single Pearson correlation that is intrinsically noisy at the 1-5 s windows real devices need; and envelope reconstruction has no native AAD rule. We argue the right object is not any single statistic but the conditional likelihood of the attended envelope given EEG, and we make this practical with NEUROTOKEN: a single network whose three heads share one EEG front-end, with a conditional flow-matching head (ATTUNEFLOW) that scores candidates by an integrated velocity-residual likelihood ratio. Two inference-time ensembles -- QUADTRACK (four complementary statistics) and ENV-FLOW (z-normalised QUADTRACK+ATTUNEFLOW) -- absorb per-statistic failure modes for free. On KU Leuven, DTU, and NJU at 5 s, ATTUNEFLOW lifts per-segment source-AAD by 9%-16% over the strongest non-generative baseline and shrinks across-subject variance by ~3x; trial-level fusion exceeds 93% on two of three datasets. In parallel reproductions we show that canonical 95-97% direction-AAD numbers collapse by 17%-45% under a strict trial-disjoint protocol, clarifying both the true ceiling and why a likelihood-based formulation is needed.