发表机构
Tampere University(坦佩雷大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究基于自然实验室和野外语音数据,利用自监督学习及多种时间池化表示建立儿童与成人定向语音分类基准,发现SSL表示性能最佳,不同池化方法互补。
AI 中文摘要
在这项工作中,我们使用来自包含自然实验室内和野外儿童定向语音(CDS)和成人定向语音(ADS)语料库的语音数据,对区分儿童定向语音与成人定向语音的分类性能进行了全面分析。我们利用自监督学习(SSL)表示,以及一系列时间池化表示(这些表示通过在高维嵌入中纳入跨通道协方差,超越了一阶和二阶统计量),为这些数据集建立了分类基准。此外,我们通过线性探针探测这些表示,以检查不同的池化方法如何捕捉韵律信息。总体而言,基于SSL的表示被证明特别有效,取得了最佳性能,而不同的池化方法为任务提供了互补优势。
英文摘要
In this work, we present a comprehensive analysis of classification performance for distinguishing child-directed speech (CDS) from adult-directed speech (ADS) using speech data from corpora containing natural in-lab and in-the-wild CDS and ADS. We establish classification benchmarks for these datasets using self-supervised learning (SSL) representations, along with a range of time-pooled representations that go beyond first- and second-order statistics by incorporating cross-channel covariances in high-dimensional embeddings. In addition, we probe these representations to examine how different pooling methods capture prosodic information using linear probes. Overall, SSL-based representations prove particularly effective, achieving the best performance, while different pooling methods offer complementary advantages for the task.
CommentsAccepted at SLT 2026