PromptEmbedder:通过双LLM软提示实现高效且可迁移的文本嵌入
PromptEmbedder: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting
- Department of Computer Science and Information Engineering, National Taiwan University(国立台湾大学计算机科学与资讯工程系)
- National Taiwan University AI Center of Research Excellence(国立台湾大学人工智能研究中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出PromptEmbedder双LLM框架,通过可微分的软提示生成将嵌入知识从特定骨干权重中解耦,在保持性能的同时降低40% GPU内存并加速3.7倍训练。
AI中文摘要:
大型语言模型(LLM)在文本嵌入方面展现出显著效果,但当前的适应方法(如LoRA)在计算效率和跨架构可迁移性方面面临重大瓶颈。每当出现新的骨干网络时,现有方法需要从头开始进行昂贵的重新训练。为了解决这个问题,我们提出了PromptEmbedder,一种新颖的双LLM框架,将嵌入知识与特定骨干权重解耦。PromptEmbedder利用一个提示LLM通过连续松弛的可微分生成过程,为冻结的嵌入LLM生成指令感知的软提示,确保对比训练期间的全梯度流动。通过将任务特定知识定位在提示LLM中,适应新架构只需重新训练一个轻量级的线性对齐矩阵。在MTEB基准上的评估表明,PromptEmbedder实现了与LoRA微调相当的性能,同时将GPU内存减少40%,训练速度提升3.7倍。我们的方法建立了一种可扩展、架构无关的范式,用于高效的基于LLM的表示学习。
英文摘要:
Large Language Models (LLMs) have demonstrated remarkable efficacy in text embedding, yet current adaptation methods like LoRA face significant bottlenecks in computational efficiency and cross-architecture transferability. Whenever a new backbone emerges, existing approaches require costly retraining from scratch. To address this, we propose PromptEmbedder, a novel dual-LLM framework that decouples embedding knowledge from specific backbone weights. PromptEmbedder utilizes a Prompting LLM to generate instruction-aware soft prompts for a frozen Embedding LLM via a differentiable generation process with continuous relaxation, ensuring full gradient flow during contrastive training. By localizing task-specific knowledge within the Prompting LLM, adapting to new architectures requires only retraining a lightweight linear alignment matrix. Evaluations on the MTEB benchmark show that PromptEmbedder achieves comparable performance with LoRA finetuning while reducing GPU memory by 40% and accelerating training by 3.7x. Our approach establishes a scalable, architecture-agnostic paradigm for efficient LLM-based representation learning.