SmallReason-ColBERT:用于推理密集型检索的超小型后期交互检索器
SmallReason-ColBERT: An Ultra-Small Late-Interaction Retriever for Reasoning Intensive Retrieval
查看机构详情
- University of Innsbruck(因斯布鲁克大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
提出32M参数的SmallReason-ColBERT后期交互检索器,通过对比预热、难负样本精炼和重要性头训练,在BRIGHT上达到21.41 nDCG@10,接近150M模型,并证明学习头优于固定IDF加权。
中文摘要 AI 辅助
推理密集型检索对于小模型而言仍然困难。紧凑的公开ColBERT通常在通用语料库上训练,在BRIGHT基准上比经过推理调优的150M+基线模型低数个nDCG@10点。然而,在边缘规模上不存在公开的推理调优ColBERT。我们提出了SmallReason-ColBERT,一个32M参数的后期交互检索器,通过三个组件缩小了这一差距:在ReasonIR-VL上进行变长对比预热,在合并的ReasonIR-HQ和BGE-Reasoner数据上进行难负样本对比精炼,以及在冻结基座之上训练的每查询词元重要性单层头。该头使用未归一化的加权MaxSim分数训练,并以其长度归一化形式进行评估。在受控重训练中,将此训练目标替换为对称归一化分数会导致损失停滞,并损失3.59个nDCG@10点。完整方案在BRIGHT上达到21.41的平均nDCG@10,比150M的Reason-ModernColBERT(22.62)低1.21,高于我们评估的所有≤33M的ColBERT。通过关于容量、初始化和分数变体的消融实验,我们进一步表明,学习到的头优于固定的IDF加权,并且简单地对学习到的门控进行阈值化是有害的。此https URL。
英文摘要
Reasoning-intensive retrieval remains difficult for small models. Compact public ColBERTs are usually trained on general-purpose corpora and underperform reasoning-tuned 150M+ baselines on BRIGHT~\cite{bright} by several nDCG@10 points. However, no public reasoning-tuned ColBERT exists at edge scale. We introduce \textbf{SmallReason-ColBERT}, a 32M late-interaction retriever that closes much of this gap with three components: a varied-length contrastive warmup on ReasonIR-VL, a hard-negative contrastive polish on merged ReasonIR-HQ and BGE-Reasoner data, and a single-layer per-query-token importance head trained on top of the frozen base. The head is trained with an un-normalised weighted MaxSim score and evaluated with its length-normalised form. In a controlled re-training, replacing this training objective with the symmetric normalised score causes the loss to stall and costs $3.59$ nDCG@10. The full recipe reaches \textbf{21.41} mean nDCG@10 on BRIGHT, within $1.21$ of the 150M Reason-ModernColBERT (22.62) and above all $\le 33$M ColBERTs we evaluate. Through ablations on capacity, initialisation, and score variants, we further show that the learned head outperforms fixed IDF weighting and that simply thresholding the learned gates is harmful. https://github.com/DataScienceUIBK/SmallReason-ColBERT