arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CS-ETS:基于混沌启发的基于桑巴的肌电到语音合成与非线性混沌损失

CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses

Sajid Fardin Dipto, Tarikul Islam Tamiti, David Vergano, Luke Baja-Ricketts, Anomadarshi Barua

arXiv 2607.18629首次发表:更新:

AI 中文总结

该研究提出CS-ETS架构用于肌电到语音合成,结合基于桑巴的编码器与LER、MSDFA两个混沌启发的损失函数,参数数量降低40.79%,新方法提升了相关指标,还减少计算量,首次实现用非线性混沌物理监督ETS,获更小高性能模型。

AI 中文摘要

我们提出了一种用于肌电到语音(ETS)合成的受混沌启发的新架构CS-ETS,它将基于桑巴的编码器与两个受混沌启发的新损失函数——李雅普诺夫指数正则化(LER)和多尺度去趋势波动分析(MSDFA)相结合。LER基于李雅普诺夫指数设计以捕捉非线性波动和对初始条件的敏感性。MSDFA利用去趋势波动分析来量化分形类、长程时间混沌相关性。CS-ETS以低40.79%的参数数量(32M对54.1M)超越了先前工作,并引入了一种新的后置声码器对齐方法,使LSD提高2.1倍,STOI提高4.7倍,SI-SDR提高1.25倍。CS-ETS在保持性能提升的同时减少了13.33%的计算量。据我们所知,首次展示了如何通过带有桑巴注意力的微妙非线性混沌物理对ETS进行监督,以实现显著更小且性能更优的模型。

英文摘要

We propose a chaos-inspired new architecture for EMG-to-Speech (ETS) synthesis called CS-ETS, which combines a Samba-based encoder with two novel chaos-inspired loss functions -- Lyapunov Exponent Regularization (LER) and Multi-Scale Detrended Fluctuation Analysis (MSDFA). LER is designed based on Lyapunov exponents to capture nonlinear fluctuations and sensitivity to initial conditions. MSDFA exploits detrended fluctuation analysis to quantify fractal-like, long-range temporal chaotic correlation. CS-ETS surpasses prior work with a 40.79\% lower parameter count (32M vs 54.1M) and introduces a new Post-Vocoder Alignment approach that improves LSD by 2.1x, STOI by 4.7x, and SI-SDR by 1.25x. CS-ETS reduces computation by 13.33\% while maintaining improved performance. To the best of our knowledge, for the first time, we show how ETS can be supervised by the subtle non-linear chaotic physics with Samba attention to achieve a significantly smaller model with superior performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑