arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27762cs.SDcs.LG

“那是什么声音?”:一种用于环境声音识别的多功能、稳健且轻量级的卷积Transformer

"What's That Sound?": A Versatile, Robust, and Lightweight Convolutional Transformer for Environment Sound Recognition

Julia Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对助听器无法识别非语音声音的问题,提出随机音频增强分层卷积Transformer(RALCT),结合MFCC与log-mel特征及CNN-Transformer架构,以约31万参数在UrbanSound8K上达94.56%准确率,并集成移动应用提升听障者安全。

中文摘要 AI 辅助

传统助听器既昂贵又使用受限,因为它并非用于检测非语音音频。我们的目标是开发一种机器学习解决方案,以提供更准确且更实惠的机制来识别周围声音,从而提高听力受损者的安全性,例如,如果汽车在行人身后鸣笛,或发生枪击,他们需要远离声源。通过向音频添加随机增强,将梅尔频率倒谱系数(MFCC)图和log-mel声谱图拼接,并在Transformer架构中纳入卷积神经网络(CNN),随机音频增强分层卷积Transformer(RALCT)模型能够高效地从多样化的音频表示中提取特征。此外,RALCT足够小巧,仅约310,000个参数,可部署到移动设备中。在UrbanSound8K数据集上的实验结果显示,RALCT所有变体的准确率持续超过93%,最高达94.56%,达到了最先进水平。为利用该技术的能力,开发了一款与模型集成的移动应用,以提供实时安全控制。因此,RALCT代表了一种稳健、轻量、实惠且多功能的深度学习工具,有助于听力受损者的导航和安全。

英文摘要

The conventional hearing aid is both costly and lim- ited in usage, as it is not intended to detect non-speech audio. Our objective is to develop a machine learning solution to provide a more accurate and affordable mechanism to identify surrounding sounds to improve the safety of the hearing impaired, i.e., if a car is honking behind pedestrians, or a gunshot is fired, and they need to move away from the source. By adding randomized augmentations to audio, concatenating a Mel-Frequency Cepstral Coefficients (MFCCs) diagram and a log-mel Spectrogram, and including Convolutional Neural Networks (CNNs) in a Trans- former architecture, the Randomized Audiomentational Layered Convolutional Transformers (RALCT) model efficiently extracts features from diversified audio representations. In addition, RALCT is small enough, with only approximately 310,000 parameters, to be deployed into mobile devices. Experimental results on the UrbanSound8K dataset resulted in an accuracy consistently over 93% for all variations of RALCT with the highest at 94.56%, reaching state-of-the-art levels. To leverage the capabilities of this technology, a mobile app is developed to be integrated with the model to provide real-time safety control. RALCT thus represents a robust, lightweight, affordable, and versatile deep learning tool to aid the navigation and safety of the hearing impaired.

发表机构

  • Northville High School(诺斯维尔高中)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑