arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

未被听到但可识别:用于乐器识别的超声波特征

Unheard but Recognizable: Ultrasonic Signatures for Musical Instrument Recognition

Izhak Kapash, Uri Rom, Ram Zamir

arXiv 2610.08850首次发表:更新:

发表机构

Tel Aviv University(特拉维夫大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究超声波成分对乐器识别的贡献,利用96-kSPS多轨语料库,通过比较全频带、低通和仅超声波输入,发现超声波信息显著提升孤立及多音识别性能。

AI 中文摘要

乐器识别是音乐信息检索(MIR)中的一项基础任务,支持转录、声源分离和音频理解。尽管算法取得了显著进展,但在识别声学上相似的乐器方面仍有改进空间,尤其是在多音混合中。我们研究了超声波成分的贡献,即高于可听范围0-20 kHz的频率。虽然根据奈奎斯特准则,这些内容大部分无法在常规的44.1/48-kSPS采样率下表示,但在高质量的96-kSPS录音中是可用的。利用一个涵盖15个类别、跨越演奏者、乐器、录音室和录音场次的96-kSPS多轨语料库,我们使用经典机器学习和深度学习分类器比较了全频带、低通和仅超声波输入。我们的结果表明,依赖于声源的超声波扩展和宽带瞬态包含乐器特定信息,显著提高了孤立和多音识别性能,并激发了对其他MIR任务进行更宽带宽研究的动机。

英文摘要

Musical-instrument recognition is a fundamental music information retrieval (MIR) task supporting transcription, source separation, and audio understanding. Despite significant algorithmic progress, there is still room for improvement in recognizing acoustically similar instruments, particularly in polyphonic mixtures. We examine the contribution of ultrasonic components, i.e., frequencies above the audible 0-20 kHz range. Although most of this content cannot be represented at conventional 44.1/48-kSPS sampling rates under the Nyquist criterion, it is available in high-quality 96-kSPS recordings. Using a 96-kSPS multitrack corpus covering 15 classes across performers, instruments, studios, and sessions, we compare full-band, low-pass, and ultrasonic-only inputs using classical machine-learning and deep-learning classifiers. Our results show that source-dependent ultrasonic extensions and broadband transients contain instrument-specific information that significantly improves isolated and polyphonic recognition and motivates wider-bandwidth studies of other MIR tasks.

Comments5 pages, 3 figures, 3 tables. Submitted to IEEE ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑