arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

内容是留存之物:基于平行话语的不变语音分词

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Laurin Wagner, Bernhard Thallinger, Miroslav Stankovic, Mario Zusag

arXiv 2607.19033首次发表:更新:

发表机构

nyra labs(nyra实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对离散语音分词器问题,提出PINT方法,通过微调SSL编码器及相关损失和增强来提取共享残差,降低条件熵,使词元保留时间基础,实验证明该方法能有效提升效果,正确不变性对高效学习很关键。

AI 中文摘要

离散语音分词器旨在将语义与声学信息分离,但像HuBERT这样的自监督学习(SSL)模型的目标会保留非语言变化:说话者身份、韵律和声道条件会渗入词元,增加熵。我们的关键见解是,当足够多的说话者在不同条件下说出相同的单词时,语言内容是唯一的共享因素。我们提出了PINT(平行不变分词),它通过跨平行话语的对齐损失和增强来微调SSL编码器,以提取这种共享残差。PINT将相同的单词合并到一致的词元序列上,大幅降低条件熵。与ASR文本不同,PINT词元保留帧级时间基础,并作为音频编解码器的直接语义目标。实验表明,与基线相比,说话者探测准确率相对降低了98.7%(从93.1%降至1.2%),ABX错误率降低了42%,语言模型困惑度降低了27 - 30%,证实了正确的不变性是高效学习的关键。

英文摘要

Discrete speech tokenizers aim to disentangle semantic from acoustic information, yet targets from self-supervised learning (SSL) models like HuBERT retain non-linguistic variation: speaker identity, prosody, and channel conditions leak into the tokens, inflating entropy. Our key insight is that when enough speakers utter the same words under varying conditions, linguistic content is the only shared factor. We propose PINT (Parallel INvariant Tokenization), which fine-tunes an SSL encoder with alignment losses across parallel utterances and augmentations to distill this shared residual. PINT collapses identical words onto consistent token sequences, drastically reducing conditional entropy. Unlike ASR text, PINT tokens preserve frame-level temporal grounding and serve as drop-in semantic targets for audio codecs. Experiments show a 98.7% relative reduction in speaker probe accuracy (93.1% to 1.2%), a 42% lower ABX error rate, and 27-30% lower LM perplexity versus baselines, confirming that the right invariance is key to efficient learning.

CommentsAccepted at Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑