基于排名感知奖励最大化的查询嵌入测试时优化
Test-Time Optimization of Query Embeddings with Ranking Aware Reward Maximization
浏览论文内容
中文总结 AI 辅助
本研究提出TTT-Embed框架,将排名奖励提炼为轻量学习向量,无需访问模型权重即可优化,在多嵌入模型和MTEB任务上显著提升测试时检索效果,且具备良好泛化性,解决了灾难性遗忘问题。
中文摘要 AI 辅助
密集检索器利用冻结编码器与预计算索引间的向量相似度对文档进行排名。尽管来自重排序器或LLM评判器的测试时排名奖励可提升结果,但现有方法会在单次查询后丢弃该信号。更新检索器权重可使奖励重复使用,但需访问参数,而闭源模型无法提供该权限,且计算成本高昂。我们提出TTT-Embed(Test-Time Tuning of Embeddings,嵌入的测试时调优),这一框架将排名奖励提炼为冻结模型输出嵌入空间内的轻量学习向量。该向量仅通过分配给检索器自身候选文档的标量排名分数进行优化,无需访问模型权重、真实标签或修改索引。单一范围参数控制奖励复用(全局、任务或查询级),可在固定奖励计算预算下实现复用性与特异性间的原则性权衡。我们证明,随着可用奖励预算扩大,最优共享范围会从全局级动态转向任务级,最终变为查询级。在5种嵌入模型和15个MTEB检索任务上评估显示,TTT-Embed使测试时检索的nDCG@10提升最多达8.36。关键在于,学习到的状态可有效泛化至未见过的查询(nDCG@10提升最多达8.57)和未见过的任务(提升最多达4.71)。此外,TTT-Embed成功解决了灾难性遗忘问题:通过完全冻结基础权重,它恢复了下降的通用能力(nDCG@10提升最多达8.00,甚至超越原始基础模型),同时保留域内专业化能力。这些结果确立了排名奖励是可复用的测试时状态,为包括闭源API在内的任何嵌入模型实现了预算高效的适配。
英文摘要
Dense retrievers rank documents using vector similarity between a frozen encoder and a precomputed index. While test-time ranking rewards from a reranker or LLM judge can improve results, existing methods discard this signal after a single query. Updating the retriever's weights makes rewards reusable, but this requires parameter access, which is unavailable for closed-source models, and is computationally prohibitive. We propose TTT-Embed (Test-Time Tuning of Embeddings), a framework that distills ranking rewards into a lightweight, learned vector within the output embedding space of a frozen model. This vector is optimized purely from scalar ranking scores assigned to the retriever's own candidate documents, requiring no access to model weights, ground-truth labels, or modifications to index. A single scope parameter controls rewards reuse (global, task, or query), enabling a principled trade-off between reusability and specificity under a fixed reward computation budget. We demonstrate that as the available reward budget scales, the optimal sharing scope shifts dynamically from global-wise to task-wise and finally to query-wise. Evaluated across five embedding models and 15 MTEB retrieval tasks, TTT-Embed improves test-time retrieval by up to +8.36 nDCG@10. Crucially, the learned states generalize effectively to unseen queries (up to +8.57 nDCG@10) and unseen tasks (up to +4.71 nDCG@10). Furthermore, TTT-Embed successfully resolves catastrophic forgetting: by leaving base weights entirely frozen, it recovers degraded general capabilities (up to +8.00 nDCG@10, even surpassing the original base model) while preserving in-domain specialization. These results establish ranking rewards as a reusable test-time state, enabling budget-efficient adaptation for any embedding model, including closed-source APIs.