arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.27017cs.IR

ProRetrieval:通过可执行程序合成学习编排混合搜索

ProRetrieval: Learning to Orchestrate Hybrid Search via Executable Program Synthesis

Chengsong You, Zhen Sun, Yunhai Hu, Junwei Zhou, Xiaoyu Cao, Binyu Li, Ziyan Zhao, Weiyao Wang, Liren Lu, Zhijie Ye, Yumo Cao, Yitao Long, Yiwei Xu, Qiyi Jiang,… 展开作者

Chengsong You, Zhen Sun, Yunhai Hu, Junwei Zhou, Xiaoyu Cao, Binyu Li, Ziyan Zhao, Weiyao Wang, Liren Lu, Zhijie Ye, Yumo Cao, Yitao Long, Yiwei Xu, Qiyi Jiang, Xuanyi Fu, Yufan Chen, Yilun Li, Rongkang Xiong, Yiran Zou, Nan Du

首次发表
浏览论文内容

中文总结 AI 辅助

ProRetrieval 将语言模型作为检索编排器,通过混合 DSL 合成可执行程序,在两个新基准上训练 Qwen3-4B,其 4B 模型在电商、邮件检索任务中优于 GPT-5.5 等主流模型及各类基线。

中文摘要 AI 辅助

现实世界的检索任务常需结合结构化约束与文本、图像的语义意图,并通过任意布尔逻辑进行组合。现有混合检索管道(如 reciprocal rank fusion 或自查询检索器)仅支持固定形式的组合;近期的强化学习检索器将语言模型训练为单一后端的查询生成器,未将异构检索路径的编排纳入其动作空间。本文提出 ProRetrieval,将语言模型重新定义为检索编排器:给定自然语言查询,它会在混合领域特定语言(DSL)中合成可执行程序,该 DSL 交错了针对结构化字段的 SQL 运算符与针对文本、图像的向量检索原语,其中 SQL 本身提供融合异构候选集的逻辑代数。我们使用 GRPO 和 DAPO 训练 Qwen3-4B,并在分层四元奖励机制下进行优化,在基于亚马逊产品和 Enron 邮件构建的两个新基准上评估性能。我们的 4B 模型在电商数据集上的 Hit@1 为 0.81,优于 GPT-5.5 的 0.69;在邮件数据集上的 Hit@1 为 0.91,优于 GPT-5.5 的 0.86,同时也优于 Claude Opus 4.7 及一系列综合的检索、大语言模型增强、结构化查询和基于图的基线方法。代码与数据可在指定链接获取。

英文摘要

Real-world retrieval often composes structured constraints with semantic intents over text and images through arbitrary Boolean logic. Existing hybrid pipelines such as reciprocal rank fusion or self-querying retrievers admit only a fixed form of composition, while recent reinforcement-learning retrievers train the language model as a query generator for a single backend, leaving the orchestration of heterogeneous retrieval paths outside its action space. We propose ProRetrieval, which recasts the language model as a retrieval orchestrator: given a natural-language query, it synthesizes an executable program in a hybrid DSL interleaving SQL operators over structured fields with vector-retrieval primitives over text and images, with SQL itself providing the logical algebra that fuses heterogeneous candidate sets. We train Qwen3-4B with GRPO and DAPO under a hierarchical four-term reward, and evaluate on two new benchmarks built from Amazon products and Enron email. Our 4B model surpasses GPT-5.5 (Hit@1 0.81 vs. 0.69 on e-commerce; 0.91 vs. 0.86 on email) and Claude Opus 4.7 and a comprehensive suite of retrieval, LLM-augmented, structured-query, and graph-based baselines. Code: https://anonymous.4open.science/r/ProRetrieval/; data: https://huggingface.co/datasets/anonymous-7219/ProRetrieval.

补充信息

↑