发表机构
Artefact Research Center; INSA Rennes, IRISA, CNRS, Université de Rennes; MICS, CentraleSupélec, Université Paris-Saclay(Artefact研究中心; 雷恩国立应用科学学院,IRISA,法国国家科学研究中心,雷恩大学; MICS,中央苏伊士高等学院,巴黎萨克雷大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对DSI框架复现难和评估不一致的问题,提出首个支持三种标识符类型的开源实现ReDSI及参数化NQ320K构建流程,实验效果优于或持平基线,并探索了模型缩减下的多种研究方向。
AI 中文摘要
可微搜索索引(DSI)框架(Tay等人,2022)已成为生成式检索的事实标准基线。然而,DSI难以复现:目前没有公开实现涵盖全部三种原始文档标识符类型(原子型、朴素型、语义型),已报告的结果差异很大,且普遍使用的NQ320K数据集是通过多种且描述不足的预处理方式从Natural Questions构建的。我们提出了ReDSI,这是首个支持全部三种标识符类型的开源DSI实现,并附带一个参数化且文档完善的NQ320K构建流程。实验上,我们取得了与先前DSI基线相比具有竞争力或更优的结果。此外,我们在模型缩减设置下进行了大量实验,涵盖检索效果、参数效率、训练方法和解码策略,为未来研究开辟了新方向。
英文摘要
The differentiable search index (DSI) framework (Tay et al., 2022) has become the de facto baseline for generative retrieval. However, DSI is hard to reproduce: no public implementation covers all three original document identifier types (atomic, naive, semantic), reported results vary widely, and the ubiquitous NQ320K dataset is built from Natural Questions through diverse and underspecified preprocessing. We introduce ReDSI, the first open-source DSI implementation supporting all three identifier types, together with a parameterizable and well-documented NQ320K construction pipeline. Experimentally, we achieve results that are competitive with or stronger than previous DSI baselines. Moreover, we conduct extensive experiments under model downscaling, covering retrieval effectiveness, parameter efficiency, training methods and decoding strategies, opening novel directions for future research.