发表机构
Alibaba Token Foundry; Stanford University(阿里巴巴Token铸造实验室; 斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过在声学预训练编码器上添加音频-描述对齐实现语义精炼,正确配对显著提升域平衡分类性能,且24层模型与领先编码器竞争力相当。
AI 中文摘要
通用音频表示必须在保留声学细节的同时,使高级概念能够跨越语音、音乐、环境声音以及不同能力的下游模型被访问。我们研究了一种声学预训练编码器的语义精炼方法,通过在BEST-RQ、重建和CTC的基础上添加音频-描述对齐。我们比较了匹配控制、打乱描述和正确配对轨迹,以区分正确对应关系与额外对比目标的作用。每个端点被冻结,并使用时间均值线性探针和序列感知LLM读取器进行评估,测试精炼后的信息是否可直接访问且对更强的模型仍然有用。在三个配对种子中,正确对齐使域平衡分类在线性探针下提高4.66个百分点,在序列感知LLM下提高2.59个百分点,且每个域均有正向变化。正确配对占线性探针增益的87%,而LLM在字幕生成中显示出最清晰的对应特定益处。密集声学目标在两种读取器下均提供互补增益。一个单独的24层延续模型在共享评估器下与领先的公共编码器保持竞争力,支持了受控研究之外的配方。
英文摘要
Universal audio representations must preserve acoustic detail while making high-level concepts accessible across speech, music, environmental sound, and downstream models of different capacities. We study semantic refinement of an acoustically pretrained encoder by adding audio-description alignment to a foundation of BEST-RQ, reconstruction, and CTC. We compare matched control, shuffled-description, and correctly paired trajectories to distinguish correct correspondence from an extra contrastive objective. Each endpoint is frozen and evaluated with a temporal-mean linear probe and a sequence-aware LLM readout, testing whether the refined information is directly accessible and remains useful to a stronger model. Across three paired seeds, correct alignment improves domain-balanced classification by 4.66 points with the linear probe and 2.59 points with the sequence-aware LLM, with positive changes in every domain. Correct pairing accounts for 87% of the linear-probe gain, while the LLM shows its clearest correspondence-specific benefit in captioning. Dense acoustic objectives provide complementary gains under both readouts. A separate 24-layer continuation remains competitive with leading public encoders under the shared evaluator, supporting the recipe beyond the controlled study.