arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自回归检索器:利用项目反馈改进通用多模态检索的查询理解

Autoregressive Retriever: Improving Query Understanding from Item Feedback for Universal Multimodal Retrieval

Jianfei Zhao, Yifan Wang, Feng Zhang, Xin Sun, Chong Feng, Zhixing Tan, Yang Luo, Boyuan Pan, Xu Kai, Yao Hu

arXiv 2610.11666首次发表:更新:

发表机构

School of Computer Science and Technology, Beijing Institute of Technology; Zhongguancun Academy; Xiaohongshu; Southeast Academy of Information Technology, Beijing Institute of Technology; Zhongguancun Laboratory(北京理工大学计算机科学与技术学院; 中关村实验室; 小红书; 北京理工大学东南信息科学研究院; 中关村实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对通用多模态检索中查询表示无法利用检索项目优化的问题,提出ARR模型,通过交替检索项目与更新查询嵌入,结合监督微调与强化学习优化,在域内和零样本基准上性能优于基线。

AI 中文摘要

通用多模态检索通常对查询进行一次编码,然后通过嵌入相似度对独立索引的项目进行排序。这种设计支持高效搜索,但即使检索到的项目可以帮助明确信息需求,查询表示也保持不变。我们引入了自回归检索器(AutoRegressive Retriever,ARR),这是一种多模态检索模型,既学习选择信息丰富的项目,又利用其内容优化后续检索。ARR 在检索项目和更新查询嵌入之间交替进行,然后使用最终嵌入对集合进行排序。监督微调通过逐步对比监督来教导编码器利用反馈。强化学习将反馈项目视为动作,并使用相关项目的最终 reciprocal rank( reciprocal rank 即 reciprocal rank,指相关项目排名的倒数)优化其选择。查询侧适配器支持针对固定项目索引进行此优化。ARR 在域内和零样本基准测试中均表现出强大的检索性能,平均优于对比基线。进一步分析表明,反馈在推理时可改善检索,且在观察任何项目之前,使用反馈进行训练也能改善初始查询嵌入。

英文摘要

Universal multimodal retrieval typically encodes a query once and ranks independently indexed items by embedding similarity. This design supports efficient search, but leaves the query representation unchanged even when retrieved items could help clarify the information need. We introduce the AutoRegressive Retriever (ARR), a multimodal retrieval model that learns both to select informative items and to use their content to refine subsequent retrieval. ARR alternates between retrieving an item and updating the query embedding, then uses the final embedding to rank the collection. Supervised fine-tuning teaches the encoder to use feedback through stepwise contrastive supervision. Reinforcement learning treats feedback items as actions and optimizes their selection using the final reciprocal rank of a relevant item. A query-side adapter enables this optimization against a fixed item index. ARR demonstrates strong retrieval performance on both in-domain and zero-shot benchmarks, outperforming the compared baselines on average. Further analyses show that feedback improves retrieval at inference time and that training with feedback also improves the initial query embedding, before any item is observed.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑