生成式检索器能在学习语义ID的同时不遗忘语言生成能力吗?
Can Generative Retrievers Learn Semantic IDs Without Forgetting How to Speak?
浏览论文内容
中文总结 AI 辅助
针对生成式检索仅检索微调会扭曲模型语言分布的问题,提出双目标框架SpeakGR及自适应版本,在保留有效检索性能的同时大幅降低语言漂移,适配交互式系统需求。
中文摘要 AI 辅助
生成式检索(GR)通过生成文档语义标识符(SID)实现端到端检索。然而,仅面向检索的微调会使预训练语言模型过度特化于SID预测,严重扭曲其自然语言分布,限制了其在需同时完成文档检索和自然语言响应生成的交互式系统中的适用性。我们提出SpeakGR,这是一个双目标框架,可在学习SID的同时保留语言生成能力。它将有监督SID学习与语言保留正则化相结合:一种同策略蒸馏目标,通过在原始文本词表上计算前向KL散度,使当前模型在学生生成的前缀上与原始模型的冻结副本对齐。我们进一步提出自适应SpeakGR,可根据观测到的语言漂移动态调整保留强度。与仅使用监督微调(SFT)相比,SpeakGR在MS MARCO数据集上将WikiText-2的前向KL散度降低81.3-93.8%,在自然问题(NQ)数据集上降低81.2-85.2%,同时在三种不同大语言模型(LLM)上均保持有效的检索性能。自适应SpeakGR在大多数设置下的检索效果优于基础SpeakGR,同时语言漂移仍远低于仅SFT的情况。
英文摘要
Generative retrieval (GR) enables end-to-end retrieval by generating document semantic identifiers (SIDs). However, retrieval-only fine-tuning can over-specialize pretrained language models to SID prediction, substantially distorting their natural-language distribution and limiting their suitability for interactive systems that must both retrieve documents and generate natural-language responses. We introduce SpeakGR, a dual-objective framework that learns SIDs while preserving language generation. It combines supervised SID learning with speak-preserving regularization: an on-policy distillation objective that aligns the current model with a frozen copy of the original model on student-generated prefixes using forward KL over the original text vocabulary. We further propose Adaptive SpeakGR, which dynamically adjusts the preservation strength based on observed language drift. Compared with SFT-only, SpeakGR reduces WikiText-2 forward KL by 81.3-93.8% on MS MARCO and 81.2-85.2% on Natural Questions (NQ) while retaining effective retrieval across three different LLMs. Adaptive SpeakGR further improves retrieval over SpeakGR in most settings while maintaining substantially lower language drift than SFT-only.
发表机构
- University of Glasgow(格拉斯哥大学)
- Brave
机构由 AI 辅助整理,请以论文原文为准。