AI 中文总结
研究提出jina-reranker-v3.5列表式重排器,通过混合注意力改进、多领域训练及三阶段自蒸馏方法,在满足效率等需求同时提升性能,如在BEIR上表现出色,在半结构化检索中有显著提升,还降低了推理延迟,并开源模型权重。
AI 中文摘要
列表式重排器是智能检索管道的判别核心,但生产部署同时需要效率、领域鲁棒性和对半结构化数据的流畅处理能力。我们提出了jina-reranker-v3.5,一个参数为0.6B的列表式重排器,它在不牺牲跨文档比较能力(这使得其前身jina-reranker-v3有效)的情况下满足了这些需求。jina-reranker-v3.5保留了jina-reranker-v3的“最后但不晚”(LBNL)交互,并沿三个轴进行了改进。它用三个滑动窗口层和两个全局层的混合调度取代了均匀全局注意力,根据LBNL读出要求将终端层固定为全局注意力。它在一个精心策划的多领域混合数据集上进行训练,涵盖法律、医学、金融、多语言和结构化检索。它通过一个三阶段自蒸馏方法传递质量,其中全注意力教师设定上限,稀疏注意力学生随后在分阶段适应协议下恢复。jina-reranker-v3.5在BEIR上达到63.20的nDCG@10,参数约少7倍的情况下与4B模型匹配,在MIRACL和RTEB上也比jina-reranker-v3有所改进。其最大的提升来自半结构化检索,nDCG@10比jina-reranker-v3提高了9.6分,领先所有可比规模的重排器。混合调度进一步将列表式推理延迟最多降低1.56倍。我们在Hugging Face上以非商业许可发布了模型权重。
英文摘要
Listwise rerankers are the discriminative core of agentic retrieval pipelines, yet production deployment demands efficiency, domain robustness, and fluency on semi-structured data at the same time. We present jina-reranker-v3.5, a 0.6B-parameter listwise reranker that meets these demands together without sacrificing the cross-document comparison that makes its predecessor jina-reranker-v3 effective. jina-reranker-v3.5 keeps the last-but-not-late (LBNL) interaction of jina-reranker-v3 and reworks it along three axes. It replaces uniform global attention with a hybrid schedule of three sliding-window layers followed by two global layers, pinning the terminal layer to global as LBNL readout requires. It trains on a curated multi-domain mixture that spans legal, medical, financial, multilingual, and structured retrieval. It transfers quality through a three-stage self-distillation recipe in which a full-attention teacher sets an upper bound that a sparse-attention student then recovers under a staged adaptation protocol. jina-reranker-v3.5 reaches 63.20 nDCG@10 on BEIR, matching a 4B model at roughly 7x fewer parameters, and improves over jina-reranker-v3 on MIRACL and RTEB as well. Its largest gains come on semi-structured retrieval, where it lifts nDCG@10 by 9.6 points over jina-reranker-v3 and leads all rerankers of comparable size. The hybrid schedule further cuts listwise inference latency by up to 1.56x. We release the model weights on Hugging Face under a non-commercial license.
Comments13 pages, 2 figures, 9 tables