发表机构
Santa Clara University; Independent Researcher(圣克拉拉大学; 独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对异常心音检测,在二分类任务中固定卷积神经网络,改变音频转图像方式,尝试常规对数梅尔频谱图、PCEN 和多分辨率版本三种方法,在数据集上测试,发现 PCEN 和多分辨率版本在准确性上优于普通对数梅尔频谱图。
AI 中文摘要
心脏病导致众多人死亡,早期检测的一种低成本方法是通过听诊器听取心音,或更好的是记录并通过模型分析。本项目是一个二分类任务,即判断某人心跳的短片段声音是否正常。我们始终使用相同的卷积神经网络(CNN),仅改变将原始音频转换为图像以供其分析的方式。我们尝试了三种方法:常规对数梅尔频谱图、PCEN(对每个频率 bin 随时间进行归一化)以及将几种不同窗口大小堆叠在一起的多分辨率版本。我们在 PhysioNet 2016 心音数据集上使用完全相同的设置(相同模型、优化器、随机种子)运行这三种方法。结果表明,三种方法在检测异常情况方面都表现良好(灵敏度约为 0.95),但在官方 PhysioNet 准确性指标上,PCEN 和多分辨率版本均优于普通对数梅尔频谱图(分别为 0.915 和 0.916 对比 0.910)。我们还运行了 Grad-CAM 以查看模型实际关注的位置,它主要聚焦于 S1 和 S2 心音所在的低频区域,这表明它学到了一些实际的东西。
英文摘要
Heart disease kills a lot of people, and one cheap way to catch it early is by listening to heart sounds with a stethoscope, or better yet, just recording them and running them through a model. This project is a binary classification task: take a short clip of someones heartbeat and decide if it sounds normal or abnormal. Instead of trying out a bunch of different models, we kept the CNN the same the whole time and just changed how we turned the raw audio into a picture for it to look at. We tried three ways of doing that: a regular logmel spectrogram, PCEN (which basically normalizes each frequency bin over time), and a multi resolution version that stacks a few different window sizes together. We ran all three on the PhysioNet 2016 heart-sound dataset with the exact same setup but same model, same optimizer, same random seed. Turns out all three do pretty well at catching abnormal cases (sensitivity around 0.95), but PCEN and multi-resolution both edge out the plain logmel on the official PhysioNet accuracy metric (0.915 and 0.916 vs. 0.910). We also ran Grad-CAM to see where the model was actually looking, and it mostly focused on the low frequencies where S1 and S2 heart sounds live, which is a good sign that it learned something real