Fretiq:通过工程频谱特征和留出的自由演奏评估实现浏览器原生电吉他弦分类
Fretiq: Browser-Native Electric Guitar String Classification via Engineered Spectral Features and Held-Out Free-Play Evaluation
浏览论文内容
中文总结 AI 辅助
研究针对单音电吉他音频弦分类难题,提出Fretiq系统,基于26维特征表示及比较训练方法,在平衡帧验证中达97.1%准确率,留出自由演奏评估总体准确率87.8%,还描述了特征提取管道及实现失败模式,且系统可在浏览器内运行。
中文摘要 AI 辅助
识别单音电吉他音频中产生给定音高的弦是一项基本的分类挑战:一个音高通常可以在不同品位位置的多根弦上产生,先前的听觉研究证实未经训练的人很难察觉到音色差异。现有使用支持向量机和频谱包络特征的方法在六弦电吉他分类中F值达到0.90,早期工作中的弦-逆频率特征F1分数高达0.72。我们提出Fretiq,一个基于浏览器的单乐器、单演奏者弦分类系统,基于26维特征表示,包括频带能量、频谱统计和13个梅尔频率倒谱系数,在322,215个平衡帧上实现了97.1%的混洗帧级验证准确率。消融研究确定MFCC是主要的准确率驱动因素。我们还引入了比较训练,通过混淆矩阵分析评估其贡献。对103,000帧的留出自由演奏评估总体准确率为87.8%。我们用Python和TypeScript描述了特征提取管道,记录了两个关键的实现失败模式。该系统完全在浏览器中运行,无需六音拾音器、指板传感器、摄像头或多麦克风设置。
英文摘要
Identifying which string produces a given pitch in monophonic electric guitar audio is a classification challenge: a single pitch can often be produced on multiple strings, with timbral differences largely imperceptible to untrained humans. We present Fretiq, a preliminary single-instrument, single-player browser-based string classification system using a 26-dimensional feature representation of frequency band energies, spectral statistics, and 13 Mel-Frequency Cepstral Coefficients. Across five seeds, a shuffled frame-level validation split yields 97.25 +/- 0.32 percent accuracy, with an ablation study identifying MFCCs as the primary accuracy driver (92.09 +/- 0.50 percent without MFCCs). We introduce Comparison Training, a data collection method recording same-pitch pairs on adjacent strings in deliberate alternation. An initial shuffled-split comparison found no net benefit but was confounded by non-comparable validation sets across conditions. A corrected matched evaluation, using an identical recording session held out from training and model selection, shows including comparison-session data improves accuracy by 25.78 +/- 1.46 percentage points; a size-matched control shows this is not explained by training-set size alone. A recording-session-held-out evaluation yields 86.53 +/- 1.23 percent accuracy, closely matching an independently collected free-play evaluation (87.8 percent), both well below the shuffled-split figure, showing shuffled validation substantially overestimates real generalization here. We describe the feature extraction pipeline in Python and TypeScript for training-inference parity and document two implementation failure modes. The system runs entirely in-browser with no specialized hardware required.