发表机构
University of California, Santa Cruz(加州大学圣克鲁兹分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出注意力中继(Attention Relay),无需训练即可将指令微调LLM的注意力权重传递给文本嵌入模型,使其具备指令感知能力,实验验证了其有效性和机制。
AI 中文摘要
通过对比学习训练的文本嵌入模型从指令配对数据中学习遵循任务指令,而经过指令微调的LLM已经知道如何遵循这些指令。我们证明,这种指令遵循能力可以在无需任何训练的情况下,从LLM迁移到基于Transformer的嵌入器。我们提出了注意力中继(Attention Relay),它将LLM产生的注意力权重传递给嵌入器自身的注意力机制。在来自Qwen3、Llama 3.1和OLMo 3家族的六个指令微调LLM,以及十个在分词器、规模和池化类型上各不相同的广泛使用的嵌入模型中,注意力中继使几乎每个组合都具备指令感知能力。将方法分解为各个部分的实验表明,LLM的注意力权重在其后期层中跟踪指令,并且这些权重主要来自指令微调。实验还表明,中继这些权重会选择文本中哪些内容重要:它使指令所要求的文本方面在嵌入中占据主导地位,或者在平均化稀释了该方面的情况下恢复该方面。
英文摘要
Text embedding models trained with contrastive learning learn to follow task instructions from instruction-paired data, while instruction-tuned LLMs already know how to follow them. We show that this instruction-following ability can carry over from an LLM to a Transformer-based embedder without any training. We propose Attention Relay, which passes the attention weights an LLM produces to the embedder's own attention. Across six instruction-tuned LLMs from the Qwen3, Llama 3.1 and OLMo 3 families and ten widely used embedding models that differ in tokenizer, size and pooling type, Attention Relay makes nearly every combination instruction-aware. Experiments that break the method down into its parts show that the LLM's attention weights track the instruction in its later layers and come largely from instruction tuning. They also show that relaying these weights selects which content in the text matters: it makes the aspect of the text that the instruction asks about dominant in the embedding, or restores that aspect where averaging had diluted it.