arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于说话人跟踪的展开式递归期望最大化神经网络

Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking

Rina Veler, Sharon Gannot

arXiv 2607.26575首次发表:更新:

AI 中文总结

该研究提出深度展开式REM网络,引入含FiLM和PE的Step Size Network动态调整递归权重,在单说话人跟踪任务中,其性能优于经典CREM基线,可应用于动态声学场景。

AI 中文摘要

我们提出了一种深度展开式REM网络,用于在轻度混响环境中鲁棒跟踪单个移动说话人。与依赖固定步长衰减策略的经典REM算法不同,所提架构通过将迭代过程展开为可微层来学习自适应更新策略。我们引入了一种Step Size Network,该网络利用FiLM和PE,基于时间上下文与收敛状态动态调整递归权重。在混响条件下的单说话人跟踪实验结果表明,所提展开式网络优于采用空间网格搜索将估计质心映射到物理位置的经典CREM基线。在单说话人跟踪任务中,所提方法实现了比CREM基线更低的RMSE,凸显了其在动态声学场景中的应用潜力。

英文摘要

We propose a deep unfolded REM network for robust tracking of a single moving speaker in mild reverberant environments. Unlike classical REM algorithms, which rely on fixed-step-size decay schedules, the proposed architecture learns an adaptive update policy by unfolding the iterative procedure into differentiable layers. We introduce a Step Size Network that leverages FiLM and PE to dynamically adjust the recursion weights based on temporal context and convergence state. Experimental results for tracking a single speaker under reverberant conditions demonstrate that the proposed unfolded network outperforms the classical CREM baseline, which employs a spatial grid search to map the estimated centroids to physical positions. In the single-speaker tracking task, the proposed method achieves a lower RMSE than the CREM baseline, highlighting its potential for dynamic acoustic scenarios.

Commentsproceedings of IWAENC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑