MyoCodec:一种用于肌电信号的流式神经编解码器
MyoCodec: A Streaming Neural Codec for Electromyography
- Signal Analysis and Interpretation Lab (SAIL), University of Southern California, USA(美国南加州大学信号分析与解释实验室(SAIL))
- Center for Language and Speech Processing, Johns Hopkins University, USA(美国约翰霍普金斯大学语言与语音处理中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
MyoCodec是一种流式神经编解码器,利用因果Transformer和残差向量量化将肌电信号编码为离散令牌,在多个下游任务中性能优异,并支持实时流式处理。
AI中文摘要:
神经编解码器将连续信号编码为紧凑的离散令牌序列,为高效传输、存储和基于令牌的序列建模提供了接口。这一范式已被现代语音和音频框架广泛采用;然而,生物信号领域仍然缺乏专门为低比特率流式传输和跨多种下游任务的泛化而设计的神经编解码器。我们提出了MyoCodec,一种专为肌电信号(EMG)设计的流式神经编解码器。受近期神经音频编解码器的启发,MyoCodec结合了因果Transformer与残差向量量化,将连续EMG信号编码为不同级别的EMG表示,范围从连续潜在特征到以50 Hz运行的离散令牌。MyoCodec在十二个公开EMG数据集上训练,在内在编解码器质量和代表性下游任务(包括打字(emg2qwerty)、手部姿态(emg2pose)、语音解码(emg2speech)和语音到EMG合成(speech2emg))中均取得了良好的性能。在这些任务中,MyoCodec在提供紧凑且因果的EMG表示的同时,展现出优于先前模型的强大性能。在流式推理期间,每个20毫秒帧仅需0.482毫秒的计算时间,从而实现实时流式传输。此外,MyoCodec提供的离散令牌表示有望支持集成到基于语言模型的方法中,为基于LLM的交互系统开辟道路,其中令牌化的EMG表示可直接被处理到此类语言或语音模型中。代码和模型权重已发布。
英文摘要:
Neural codecs encode continuous signals into compact sequences of discrete tokens, providing an interface for efficient transmission, storage, and token-based sequence modeling. This paradigm has been widely adopted in modern speech and audio frameworks; however, the biosignal domain still lacks a neural codec designed specifically for low-bitrate streaming and generalization across diverse downstream tasks. We present MyoCodec, a streaming neural codec designed for electromyography (EMG). Inspired by recent neural audio codecs, MyoCodec combines causal Transformers with residual vector quantization to encode continuous EMG signals into different levels of EMG representations spanning from continuous latent features to discrete tokens operating at 50 Hz. Trained on twelve public EMG datasets, MyoCodec achieves favorable performance in both intrinsic codec quality and representative downstream tasks, including typing (emg2qwerty), hand-pose (emg2pose), speech decoding (emg2speech), and speech-to-EMG synthesis (speech2emg). Across these tasks, MyoCodec exhibits strong performance against prior models while providing a compact and causal EMG representation. During streaming inference, it requires compute time of only 0.482 ms for each 20 ms frame, enabling real-time streaming. Also, the discrete token representation provided by MyoCodec has the potential to support integration into language-model based approaches, creating a path toward LLM-based interactive systems, where tokenized EMG representations are directly processed into such language or speech models. Code and model weights are released.