arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Indic-CLAP:通过跨语言蒸馏实现印度语言的文本-音频理解

Indic-CLAP: Text--Audio Understanding in Indic Languages via Cross-Lingual Distillation

Sanat Kumar Agrawal, Srikanth Raj Chetupalli

arXiv 2610.08846首次发表:更新:

发表机构

Indian Institute of Technology Bombay(印度孟买理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Indic-CLAP通过跨语言蒸馏将CLAP扩展到九种印度语言,仅训练文本编码器并冻结音频编码器,实现了跨模态检索和零样本分类,性能趋势由模态间隙分析解释。

AI 中文摘要

对比语言-音频预训练(CLAP)学习一个联合的文本-音频表示空间,从而能够从自然语言描述中进行零样本音频理解。然而,其仅支持英文的文本编码器将CLAP限制在英文表达的任务和数据集上。我们提出了Indic-CLAP,一个多语言文本-音频模型,将CLAP扩展到九种印度语言,这些语言有超过十亿人使用。利用CLAP的双编码器结构,我们提出仅通过蒸馏训练一个印度文本编码器,同时保持音频编码器冻结,使用机器翻译的AudioCaps字幕。我们比较了从CLAP的文本编码器、音频编码器以及结合两者的新型混合目标进行蒸馏的效果。我们的评估表明,Indic-CLAP编码器继承了CLAP教师的跨模态对齐能力,这通过跨模态检索和零样本分类任务的性能得以证明。此外,我们使用模态间隙分析研究了跨模态对齐,这解释了性能趋势。

英文摘要

Contrastive Language--Audio Pretraining (CLAP) learns a joint text--audio representation space, which enables zero-shot audio understanding from natural-language descriptions. However, its English-only text encoder restricts CLAP to tasks and datasets expressed in English. We present Indic-CLAP, a multilingual text--audio model that extends CLAP to nine Indic languages spoken by over a billion people. Exploiting CLAP's dual-encoder structure, we propose to train only an Indic text encoder by distillation while keeping the audio encoder frozen, using machine-translated AudioCaps captions. We compare distillation from CLAP's text encoder, audio encoder, and a novel hybrid objective combining both. Our evaluation shows that the Indic-CLAP encoder inherits the cross-modal alignment from the CLAP teacher, as evidenced by the performance on cross-modal retrieval and zero-shot classification tasks. Further, we study the cross-modal alignment using modality gap analysis, which explains the performance trends.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑