发表机构
ABV-Indian Institute of Information Technology and Management(ABV-印度信息技术与管理学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出19维物理信息描述符集MADS,在ESC-10、ESC-50、MSoS数据集上,其分类性能优于26维MFCC基线和38维频谱摘要基线,可作为音频建模的基础描述符层。
AI 中文摘要
主流音频分类流程在深度建模前要么依赖紧凑的手工设计摘要,要么依赖固定的时频前端(如对数梅尔表示)。这些表示虽取得了较高的成功,但未明确揭示潜在发声事件的物理动力学特性。我们提出MADS(Multi-view Acoustic Descriptor Set,多视图声学描述符集),这是一个紧凑的19维物理信息描述符集,旨在捕捉音频信号中互补的频谱、时间、机械和随机结构。MADS并非仅将声音视为频谱模式,而是在统一的多视图表示中编码与激励、阻尼、周期性、脉冲性及结构一致性相关的属性。我们在ESC-10、ESC-50和MSoS数据集上使用标准经典机器学习模型评估MADS,并将其与两种传统手工基线对比:一种是紧凑的26维基于MFCC的基线,另一种是扩展的38维频谱摘要基线。在ESC-10和ESC-50上,MADS整体达到最强峰值结果,分别为81.00%和52.78%,同时其维度约为38维基线的一半;在MSoS上,MADS再次实现最强的顶级性能,达到67.48%。这些结果表明,MADS不仅是具有竞争力的独立描述符集,更是未来面向帧级、与深度学习兼容的音频建模的更广泛声学基础表示方案的基础描述符层。
英文摘要
Dominant audio classification pipelines rely either on compact handcrafted summaries or on fixed time-frequency frontends such as log-mel representations prior to deep modeling. While highly successful, these representations do not explicitly expose the physical dynamics of the underlying sound-generating event. We introduce MADS (Multi-view Acoustic Descriptor Set), a compact 19-dimensional physics-informed descriptor set de- signed to capture complementary spectral, temporal, mechanical, and stochastic structure in audio signals. Rather than treating sound only as a spectral pattern, MADS encodes properties related to excitation, damping, periodicity, impulsiveness, and structural consistency within a unified multi-view representation. We evaluate MADS using standard classical machine learning models on ESC-10, ESC-50, and MSoS, and compare it against two conventional handcrafted baselines: a compact 26D MFCC- based baseline and an expanded 38D spectral-summary baseline. Across ESC-10 and ESC-50, MADS achieves the strongest peak results overall, reaching 81.00% and 52.78%, respectively, while using roughly half the dimensionality of the 38D baseline. On MSoS, MADS again delivers the strongest top-end performance, reaching 67.48%. These results establish MADS not merely as a competitive standalone descriptor set, but as the foundational descriptor layer of a broader acoustically grounded representation program for future frame-level and deep-learning-compatible audio modeling.