发表机构
MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
RECAST通过自适应证据路由将证据构建视为顺序决策过程,利用RouterLM和CompilerLM主动推导证据,在六个基准上平均成功率75.6%,比最强基线高15.9%。
AI 中文摘要
大型语言模型越来越多地应用于基于长且异构的信息源的任务。传统的检索增强生成(RAG)依赖于固定的基于相似性的检索,而智能体变体则调整查询和使用工具,但仍在很大程度上以检索为中心。然而,在许多任务中,解决方案所需的证据并不明确存在于任何单个源项中。相反,它必须通过跨多个源项的过滤、聚合或计算来推导。在这项工作中,我们引入了RECAST(通过计算、访问和合成工具路由证据),这是一个学习框架,将证据构建表述为对异构检索和计算操作的顺序决策过程,允许证据被主动推导而非仅仅检索。一个轻量级的RouterLM迭代地选择和制定原始操作,或为冻结的CompilerLM指定定制操作以翻译成可执行代码。一旦它判断证据充分,RouterLM将接受的证据传递给冻结的AnswerLM以产生最终解决方案。我们使用监督微调(SFT)后跟组相对策略优化(GRPO)来训练RouterLM。在六个异构基准家族中,RECAST实现了75.6%的平均成功率,比最强的大型模型基线高出15.9%。此外,训练使Qwen3.5-9B RouterLM优于无需训练的Gemini 3.5 Flash RouterLM 5.0%。在三个保留的基准上,RECAST平均比最强基线提高了15.0%,展示了跨任务和异构源表示的强大零样本泛化能力。
英文摘要
Large language models are increasingly applied to tasks grounded in long, heterogeneous information sources. Conventional Retrieval-Augmented Generation (RAG) relies on fixed similarity-based retrieval, while agentic variants adapt queries and tool use but remain largely retrieval-centric. However, in many tasks, the evidence required for a solution is not explicitly present in any single source item. Instead, it must be derived through filtering, aggregation, or computation across multiple source items. In this work, we introduce RECAST (Routing Evidence through Computation, Access, and Synthesized Tools), a learned framework that formulates evidence construction as a sequential decision process over heterogeneous retrieval and computation operations, allowing evidence to be actively derived rather than merely retrieved. A lightweight RouterLM iteratively selects and formulates primitive operations or specifies customized operations for a frozen CompilerLM to translate into executable code. Once it judges the evidence sufficient, RouterLM passes the accepted evidence to a frozen AnswerLM to produce the final solution. We train RouterLM with supervised fine-tuning (SFT) followed by group relative policy optimization (GRPO). Across six heterogeneous benchmark families, RECAST achieves a mean success rate of 75.6%, outperforming the strongest large-model baseline by 15.9%. Moreover, training enables the Qwen3.5-9B RouterLM to outperform a training-free Gemini 3.5 Flash RouterLM by 5.0%. On three held-out benchmarks, RECAST improves over the strongest baseline by 15.0% on average, demonstrating strong zero-shot generalization across tasks and heterogeneous source representations.
Comments35 pages, 2 figures, 13 tables