可变速率谐波-打击乐时间尺度修改及其Python实时播放
Variable-Rate Harmonic-Percussive Time-Scale Modification with Real-Time Playback in Python
- Harvey Mudd College(哈维穆德学院)
- University of California, Los Angeles(加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出一种Python实现的谐波-打击乐时间尺度修改方法,支持可变速率实时播放,通过预计算表优化减少约一半运行时间,且感知质量不变。
AI中文摘要:
时间尺度修改(TSM)已有多个开源实现,但这些实现几乎专为离线使用而设计,即录音以固定速率处理并提前写出。自动音乐伴奏等应用需要一种不同的设置,我们称之为可变速率播放:待拉伸的录音事先已知,但播放速率未知,且必须随现场演奏者连续变化。少数实时生成输出的实现是用C++编写的,并针对速度而非易修改性、实验性以及与主要基于Python的研究生态系统的集成性进行了优化。本文描述了一种用于可变速率播放设置的广泛使用的谐波-打击乐TSM方法的Python实现。谐波-打击乐分离作为预处理步骤在已知输入录音上离线执行,而合成和播放则以可每帧变化的时间尺度因子实时进行。我们进一步提出了一系列变体,通过用预计算表的查找替换相位声码器的分析阶段FFT和瞬时频率计算来减少运行时间。由24名参与者(1114个成对评级)进行的主观听力测试表明,一旦预计算的跳跃大小足够小,这些近似在感知上与精确实现无法区分,同时将总运行时间减少约一半。我们描述了预计算、运行时间、内存和感知质量之间的权衡,以指导算法选择,并将我们的实现作为开源软件包发布。
英文摘要:
Time-scale modification (TSM) has a number of open-source implementations, but these are designed almost exclusively for offline use, in which a recording is processed at a fixed rate and written out ahead of time. Applications such as automatic musical accompaniment require a different setting, which we call variable-rate playback: the recording to be stretched is known in advance, but the playback rate is not, and must change continuously in response to a live performer. The few implementations that generate output in real time are written in C++ and optimized for speed rather than for ease of modification, experimentation, and integration with the primarily Python-based research ecosystem. This paper describes a Python implementation of the widely used harmonic-percussive TSM method for the variable-rate playback setting. Harmonic-percussive separation is performed offline as a preprocessing step on the known input recording, while synthesis and playback are carried out in real time with a time-scale factor that may change at every frame. We further propose a family of variants that reduce runtime by replacing the phase vocoder's analysis-stage FFT and instantaneous frequency calculations with lookups into precomputed tables. Subjective listening tests with 24 participants (1114 pairwise ratings) show that these approximations become perceptually indistinguishable from the exact implementation once the precomputed hop size is sufficiently small, while reducing total runtime by roughly half. We characterize the resulting tradeoffs among precomputation, runtime, memory, and perceptual quality to guide algorithm selection, and we release our implementation as an open-source package.