发表机构
Indian Institute of Technology Gandhinagar(印度甘地讷格尔印度理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多语言文本嵌入模型适配的任务差异问题,提出TCFM框架,结合流匹配与教师引导表示保留等方法,在Indic基准上实现最优性能并将发布代码与数据集。
AI 中文摘要
多语言文本嵌入模型通常采用单一训练目标在各类任务中进行适配,然而不同任务需要本质上不同的优化策略。我们提出任务条件流匹配(Task-Conditional Flow Matching, TCFM),这是一种多语言嵌入适配框架,它选择性地将流匹配应用于翻译任务,同时采用与学习动态更匹配的目标优化检索、分类和对分类任务。该框架进一步结合教师引导的表示保留与三阶段课程学习以实现稳定适配。在Indic大规模文本嵌入基准(Indic Massive Text Embedding Benchmark)上评估,TCFM取得了新的 state-of-the-art,在各类多语言任务中持续提升嵌入质量,并能在不同嵌入模型家族间泛化。论文录用后我们将公开发布代码库和数据集。
英文摘要
Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (TCFM), a multilingual embedding adaptation framework that selectively applies Flow Matching to translation tasks while optimizing retrieval, classification, and pair-classification tasks with objectives better aligned to their learning dynamics. TCFM further combines teacher-guided representation preservation with a three-stage curriculum to enable stable adaptation. Evaluated on the Indic Massive Text Embedding Benchmark, TCFM establishes a new state-of-the-art, consistently improving embedding quality across a diverse set of multilingual tasks and generalizing across embedding model families. We will publicly release the codebase and datasets upon acceptance of the paper.