arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20563cs.IR

推理质量至关重要:对抗基于LLM的嵌入学习中的推理崩溃

Reasoning Quality Matters: Combating Reasoning Collapse in LLM-based Embedding Learning

Zihan Gong, Xiaohan Ye, Jiangchao Yao, Jinsong Lan, Xiaoyong Zhu, Xu Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM嵌入学习中的推理崩溃问题,提出CoFree两阶段框架,通过参考引导微调和双重奖励强化学习,提升检索性能,在22个数据集上平均提升2.8 nDCG@10。

中文摘要 AI 辅助

大型语言模型(LLMs)近期在生成用于检索的上下文丰富的文本嵌入方面展现出强大潜力。大多数现有方法要么将嵌入学习视为被动特征提取,要么通过指令遵循利用LLM推理来优化嵌入。然而,针对嵌入目标的专门化可能抑制有用的推理生成或产生与检索无关的文本。我们将这两种退化形式称为推理崩溃。为解决此问题,我们提出CoFree(无崩溃推理嵌入),一个两阶段框架,逐步将LLM推理整合到查询和文档嵌入优化中,同时保持推理质量。在第一阶段,CoFree应用参考引导的监督微调来恢复基础嵌入模型的推理能力并保留其表示强度。在第二阶段,我们引入双重奖励,即面向嵌入的奖励和面向推理的奖励,以在强化学习中确保对嵌入目标相关性的细粒度推理。这种端点耦合优化将嵌入学习从静态对齐转变为高质量推理引导的检索搜索过程。大量实验证明了CoFree的有效性,其中CoFree-4B在来自MTEB和BRIGHT的22个数据集上,相较于Qwen3-Embedding-4B,平均绝对改进达2.8个nDCG@10点。在真实检索系统中的在线实验也显示出一致的增益。代码、RTED和模型检查点将公开发布。

英文摘要

Large Language Models (LLMs) have recently shown strong potential for producing context-rich text embeddings for retrieval. Most existing methods either treat embedding learning as passive feature extraction or exploit LLM reasoning through instruction following for better embedding optimization. However, specialization toward embedding objectives can suppress useful reasoning generation or produce retrieval-irrelevant text. We refer to these two forms of degradation as reasoning collapse. To address this issue, we propose CoFree (Collapse-Free Reasoning Embedding), a two-stage framework that progressively integrates LLM reasoning into query and document embedding optimization while preserving reasoning quality. At the first stage, CoFree applies reference-guided supervised fine-tuning to restore the reasoning ability and retain representational strength of the foundation embedding model. At the second stage, we introduce dual rewards, an embedding-oriented reward and a reasoning-oriented reward, to guarantee fine-grained reasoning of the relevance toward the embedding goal in reinforcement learning. This endpoint-coupled optimization transforms embedding learning from static alignment into a high-quality reasoning-guided search process for retrieval. Extensive experiments demonstrate the effectiveness of CoFree, with CoFree-4B achieving an average absolute improvement of 2.8 nDCG@10 points over Qwen3-Embedding-4B across 22 datasets from MTEB and BRIGHT. Online experiments in a real-world retrieval system further show consistent gains. Code, RTED, and model checkpoints will be made publicly available.

补充信息

↑