发表机构
National Chi Nan University; National Taiwan Normal University(国立暨南国际大学; 国立台湾师范大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对时频掩蔽分离器的两个限制,提出SEAL,采用零和加性残差重建和基于声学与步间证据的专家路由,在EchoSet上以更少参数和计算量达到更优或相当的分离性能。
AI 中文摘要
通过共享单元对混合信号进行掩蔽和细化的紧凑时频分离器面临两个限制。首先,有界乘法掩蔽仅缩放混合单元,因此当重叠分量相互抵消时,估计值保持较小。其次,共享单元在每一步对每个时频标记应用相同的权重,因此扩大它会在所有地方增加计算量。我们提出SEAL(稀疏专家路由与加性潜在重建)来解决这两个问题。对于重建,一个由局部混合幅度约束的零和加性残差使得估计值在分量抵消处非零,同时仍能总和为混合信号。对于路由,一个由声学和步间证据构建的查询将每个标记发送到六个残差专家之一,并且范数上限防止步进提示覆盖清晰的声学证据。在EchoSet上,SEAL(小)以28%更少的参数和2.9倍更少的MACs超过TIGER(小)0.31 dB SI-SDRi,而SEAL(大)在3.1倍更少的MACs下与TIGER(大)的SI-SDRi差距在0.07 dB以内。
英文摘要
Compact time-frequency separators that mask the mixture and refine through a shared cell face two limits. First, a bounded multiplicative mask only scales a mixture bin, so where overlapping components cancel, the estimate stays small. Second, a shared cell applies the same weights to every time-frequency token at every step, so enlarging it adds compute everywhere. We present SEAL (Sparse Expert routing with Additive Latent reconstruction) to address both. For reconstruction, a zero-sum additive residual bounded by the local mixture amplitude lets estimates be nonzero where components cancel yet still sum to the mixture. For routing, a query built from acoustic and inter-step evidence sends each token to one of six residual experts, and a norm cap keeps the step cue from overriding clear acoustic evidence. On EchoSet, SEAL (small) surpasses TIGER (small) by 0.31 dB SI-SDRi with 28% fewer parameters and 2.9 times fewer MACs, and SEAL (large) is within 0.07 dB SI-SDRi of TIGER (large) at 3.1 times fewer MACs.
CommentsSubmitted to ICASSP 2027