AI 中文总结
本研究提出面向土耳其语的晚交互多向量检索模型MoganColBERT-TR,以MoganBERT-TR编码器为起点经单轮蒸馏训练,在5个土耳其语BEIR零样本数据集上取得良好性能,位列对比模型第二。
AI 中文摘要
我们此前报告了一个从头训练的土耳其语ModernBERT编码器(MoganBERT-TR)以及基于该编码器构建的单向量嵌入模型(MoganBERT-embed)。本研究推出该系列的第三个模型:MoganColBERT-TR,这是一个多向量检索模型,它不将查询或文档压缩为单个向量,而是通过768→128的投影以词元级别表示,并通过MaxSim晚交互进行打分。该模型并非从头训练:以嵌入模型的编码器为起点,通过单轮蒸馏阶段适配到ColBERT目标函数。训练数据来自两个来源:一是从我们自身字符域预训练语料库的句子边界处切分的标题-段落对,二是两个土耳其语问答检索集;这些数据从交叉编码器教师模型(bge-reranker-v2-m3)的软得分中蒸馏而来,该教师模型针对1个正样本和7个挖掘出的负样本进行训练。我们表明,在难负样本挖掘中,仅基于排名的跳过是不够的,必须结合组掩码和余弦上限。评估使用TurkColBERT的官方流程进行,TurkColBERT是为土耳其语晚交互检索构建的基准(采用PLAID索引、精确MaxSim),在5个土耳其语BEIR数据集上开展;这些数据集均未出现在我们的训练池中,因此所有5个结果均为干净的零样本结果。MoganColBERT-TR有1.489亿个参数,在5个数据集上的平均总得分达37.36(nDCG@100为35.53,nDCG@10为31.81),在对比的5个模型中排名第二:它在5个数据集中的4个上优于规模大一倍的ColmmBERT-base-TR,总得分提升了+3.05,且比基准中最大的模型提升了+12.30。与领先模型(mLateOn)的差距集中在ArguAna-TR数据集上,该数据集的查询长度远超其他数据集。
英文摘要
We previously reported a ModernBERT encoder trained from scratch for Turkish (MoganBERT-TR) and a single-vector embedding model built on top of it (MoganBERT-embed). This work introduces the third model in that lineage: MoganColBERT-TR, a multi-vector retrieval model that, instead of compressing a query or a document into a single vector, represents it at the token level through a 768->128 projection and scores it with MaxSim late interaction. The model is not trained from scratch: the embedding model's encoder is taken as the starting point and adapted to the ColBERT objective with a single-epoch distillation phase. Training data is produced from two sources - title-to-passage pairs carved out of our own pretraining corpus in the character domain and at sentence boundaries, and two Turkish question-based retrieval sets - and is distilled from the soft scores of a cross-encoder teacher (bge-reranker-v2-m3) over one positive and seven mined negatives. We show that in hard negative mining, rank-based skipping alone is insufficient and must be combined with a group mask and a cosine ceiling. Evaluation is carried out with the official pipeline of TurkColBERT, a benchmark built for Turkish late-interaction retrieval (PLAID index, exact MaxSim), on five Turkish BEIR datasets; none of them appears in our training pool, so all five results are clean zero-shot. With 148.9M parameters, MoganColBERT-TR reaches an overall score of 37.36 (35.53 nDCG@100, 31.81 nDCG@10) averaged over the five datasets and finishes second among the five models compared: it outperforms the twice-as-large ColmmBERT-base-TR on four of five datasets and by +3.05 overall, and the benchmark's largest model by +12.30. The gap to the leading model (mLateOn) is concentrated on ArguAna-TR, the dataset with by far the longest queries.
Comments12 pages, 6 tables. Model weights: https://huggingface.co/moganai