arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29780cs.SD

连续音频编码器中潜在维度与帧率的联合分析

Joint Analysis of Latent Dimensionality and Frame Rate in Continuous Audio Encoders

  • KAIST AI(韩国科学技术院人工智能学院)
  • Seoul National University(首尔大学)
  • NAVER Cloud(NAVER云)

机构由 AI 辅助整理,请以论文原文为准。

Kyudan Jung, Sehyun Lee, Song-ha Jo, Jaegul Choo, Sanghyuk Choi

中文总结 AI 辅助

本研究联合分析连续音频编码器的潜在宽度与帧率,发现二者在下游任务中存在交互作用,且表征组织方式比重建保真度更关键。

中文摘要 AI 辅助

连续音频编码器通过潜在宽度和帧率沿特征轴和时间轴压缩音频,但二者对下游性能的联合影响仍不清楚。我们训练了十六个编码器,涵盖四种宽度和四种帧率,并使用匹配的训练协议,配备下游适配器和探针。尽管较大宽度通常能改善重建效果,但自动语音识别(ASR)和口语问答(SQA)在较高帧率下更偏好中等宽度,且在更强的时间压缩下,最佳观测宽度向更大值偏移。冻结模型的PCA干预揭示了重建与识别敏感性的差异:在高帧率512维编码器中,移除后半部分分量会显著降低ASR性能,而重建损失相对较小;相比之下,1024维编码器在很大程度上同时保留了重建和识别性能。然而,投影后的1024维模型在12.5Hz帧率下的ASR性能不如未修改的较窄模型。这些发现揭示了宽度与帧率在下游效用中的交互作用,并表明训练过程中表征的组织方式比重建保真度和可压缩性更为重要。

英文摘要

Continuous audio encoders compress audio along feature and time axes through latent width and frame rate, but their joint effect on downstream performance remains unclear. We train sixteen encoders spanning four widths and four frame rates, with downstream adapters and probes, using matched training protocols. Despite generally improved reconstruction at larger widths, automatic speech recognition (ASR) and spoken question answering (SQA) favor moderate widths at higher rates, with the best observed widths shifting toward larger values under stronger temporal compression. Frozen-model PCA interventions reveal distinct reconstruction and recognition sensitivities: removing the trailing half of the components substantially degrades ASR in high-rate 512-dimensional encoders with comparatively small reconstruction penalties, whereas 1024-dimensional encoders largely preserve both. Yet the projected 1024-dimensional model underperforms unmodified narrower models on ASR at 12.5Hz. These findings identify a width--rate interaction in downstream utility and suggest that how representations are organized during training matters beyond reconstruction fidelity and compressibility.

补充信息

↑