将非最大概率映射到GMM分量对S-JEPA编码器表示是否重要?
Cluster Assignments in Soft Targets Shape Speech Representations: Evidence from S-JEPA
浏览论文内容
中文总结 AI 辅助
该研究针对S-JEPA,通过两个对照实验发现,软目标的非最大概率到GMM分量的映射对编码器表示有显著影响,而非仅概率结构决定表示。
中文摘要 AI 辅助
S-JEPA使用软高斯混合模型(GMM)后验而非硬聚类标签来保留不确定性。目前尚不清楚仅概率值是否足够,还是非最大概率分配给哪个GMM分量也很重要。我们通过两个匹配对照组对此进行测试:FIXED-RANDPERM保留top-1分量及其概率,以及非最大概率值的多重集,但针对每个物理帧固定映射重新分配这些非最大值;UNIFORM-TAIL保留top-1分量、其概率及总非最大质量,但将该质量均匀分布。在三个独立随机种子下,REAL SOFT在两个冻结编码器读出任务上均优于两个对照组,它能更好地恢复原始GMM尾部,且在控制当前帧完整频谱后,更易获取短时间尺度上的频谱动力学。在两个曝光实验中,保留原始映射的帧数越多,两个读出任务的整体性能均提升。我们还定性跟踪了切换到在线GMM后的一条第2阶段轨迹。这些结果表明,软目标的数值概率结构并不完全决定学习到的编码器表示,非最大概率到GMM分量的映射也很重要。
英文摘要
Cluster-based prediction is widely used in self-supervised speech learning. A soft target preserves a distribution over clusters rather than a single label. This distribution specifies both the probability values and which clusters receive them. Comparisons between soft targets and hard labels do not separate the contributions of these two aspects to the learned representation. We study this in S-JEPA, a recent high-performing self-supervised speech model trained with soft Gaussian mixture model (GMM) targets. We compare its original targets with counterfactual targets that preserve the most likely cluster and all probability values but change which remaining clusters receive the other probabilities. Across three training seeds, the original soft distribution is recovered more accurately from Encoders trained with the original than counterfactual targets. Because this could reflect target matching alone, we also test low-level acoustic and phonetic information. Both are more accessible from Encoders trained with the original targets. This suggests that cluster assignments affect acoustic and phonetic properties of the learned representation, not just recovery of the training target.