arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多智能体听觉场景分析:通过多波束形成语音质量反馈提高定位速度与鲁棒性

Multi-agent Auditory Scene Analysis: Improved Localization Speed and Robustness by Multi-beamformed Speech Quality Feedback

Caleb Rascon

arXiv 2610.00538首次发表:更新:

发表机构

Instituto de Investigaciones en Matematicas Aplicadas y en Sistemas, Universidad Nacional Autonoma de Mexico(墨西哥国立自治大学应用数学与系统研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多智能体听觉场景分析优化速度慢的问题,提出基于多波束形成语音质量反馈的新优化机制,利用多位置质量估计集简化搜索空间,显著提升定位速度、准确性与鲁棒性,且复杂度更低。

AI 中文摘要

实时听觉场景分析器(ASA)旨在完成对给定声学环境中存在的声源进行定位、分离和分类的任务。近期,已有研究尝试将ASA建模为多智能体系统,其中每个智能体执行上述任务之一,并将其结果传达给其他同级智能体。这些通信路径被用作反馈回路,以在全局层面修正局部错误,在降低局部复杂性的同时提供鲁棒性。该方法的一个优势示例是通过实时修正感兴趣语音源的估计位置来优化语音质量。然而,其优化速度已被证明相当缓慢。一个可能的原因是它仅依赖一系列单一质量估计(由无参考质量估计模型提供),这些估计在相邻窗口间变化显著,导致搜索空间难以优化。在本工作中,提出了一种新的优化机制,该机制转而依赖在一系列位置上的多组质量估计,从而提供更清晰的搜索空间视图并简化优化过程。所提出的ASA现在具有显著更短的优化时间,在真实声学场景中评估以修正更高级别的定位误差时更加准确且更稳定,同时其复杂度低于先前工作。唯一的权衡是质量估计智能体的响应时间有所增加,但完整的ASA仍能实时运行。本工作展示的性能再次证明了将ASA建模为多智能体系统的优势。

英文摘要

A real-time auditory scene analyzer (ASA) aims to carry out the tasks of locating, separating and classifying the sound sources present in a given acoustic environment. Recently, an effort has been made into modelling an ASA as a multi-agent system, with each one of its agents performing one of the aforementioned tasks and communicating their results to the rest of their peer agents. These communication routes are used as feedback loops to fix local errors at a global level, providing robustness while reducing local complexity. An example of the benefits of this approach is the optimization of speech quality by correcting in real-time the estimated location of the speech source of interest. However, their optimization speed has been shown to be considerably slow. One possible reason is that it solely relies on a series of single quality estimations (provided by a reference-free quality estimator model) that vary considerably from one window to the next, which results in a difficult search space to optimize. In this work, a new optimization mechanism is proposed that instead relies on a series of sets of quality estimations over a range of locations, providing a clearer view of the search space, simplifying its optimization. The proposed ASA now has a considerably smaller optimization time, is more accurate, and is more stable when being evaluated in real-life acoustic scenarios to correct higher levels of localization errors, all while being less complex than previous efforts. The only trade-off is that there is an increase in the response time of the quality estimation agent, but the complete ASA is still able to run in real-time. The performance shown in this work again shows the benefits of modelling an ASA as a multi-agent system.

CommentsSubmitted to Autonomous Agents and Multi-Agent Systems

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑