发表机构
TurkuNLP; University of Turku; Ellis Institute Finland(图尔库NLP; 图尔库大学; 芬兰ELLIS研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究嵌入模型在非对称检索中遵循指令的机制,发现查询侧干扰项导致模型失败,通过加入此类干扰项微调可显著提升性能且对其他任务影响极小。
AI 中文摘要
提示式嵌入模型近年来受到越来越多的关注,尤其是在检索领域,其中详细的检索指令作为检索提示的一部分被提供。多个新数据集和研究已对这一设置进行了考察,结果表明当前的嵌入模型往往难以可靠地遵循此类指令。本文研究了在非对称检索任务中,指令实际上如何影响检索查询的表示。我们证明,当评估中包含查询侧干扰项时,模型甚至可能无法遵循简单的任务指令。我们假设这一行为是由当前嵌入模型的训练设置及其评估方式所驱动的,并表明在加入查询侧干扰项进行微调后,模型性能得到显著提升,而对其他任务的影响微乎其微。
英文摘要
Prompted embedding models have recently received increasing attention, particularly for retrieval, where detailed retrieval instructions are provided as part of the retrieval prompt. Several new datasets and studies have examined this setting, showing that the current embedding models often struggle to follow such instructions reliably. In this paper, we study the mechanism of how instructions actually affect the representations of retrieval queries in asymmetric retrieval tasks. We show that models can fail to follow even simple task instructions when query-side distractors are included in the evaluation. We hypothesize that this behavior is driven by the training setup of current embedding models and their evaluation, and show that fine-tuning with added query-side distractors leads to substantial improvements, with minimal effect on other tasks.