AI 中文总结
研究针对生成式大语言模型在复杂查询时的事实幻觉及重排延迟问题,提出TALRanker框架,将相关性评分形式化为马尔可夫决策过程,经两阶段训练优化,在检索基准上性能领先,兼顾吞吐量与准确性。
AI 中文摘要
生成式大语言模型(LLMs)变革了信息检索,但面对超出认知边界的复杂查询时,其严格的参数性质常导致严重的事实幻觉。外部工具调用可缓解此问题,但重排时对每个文档都不加区分地调用搜索工具会带来高昂的延迟开销,造成难以解决的准确性 - 效率困境。为应对这一挑战,我们提出了TALRanker,一个将逐点相关性评分形式化为智能体马尔可夫决策过程的新颖框架。我们通过两阶段训练范式对其进行优化。初始热身阶段利用保留语言的混合损失来防止原生生成能力的灾难性遗忘。随后,强化学习中配备的非对称成本感知奖励迫使策略在自信时自主绕过工具以实现最大效率,而在不确定时选择性地检索外部证据以避免严重的幻觉惩罚。广泛评估表明,TALRanker在标准和推理密集型检索基准上均实现了领先性能,在吞吐量上与逐点重排器匹配,同时优于参数繁重的推理模型。
英文摘要
Generative Large Language Models (LLMs) have revolutionized information retrieval, yet their strictly parametric nature frequently leads to severe factual hallucinations when confronted with complex queries beyond their epistemic boundaries. While external tool-calling can mitigate this, indiscriminately invoking search tools for every document during reranking incurs prohibitive latency overheads, creating an intractable accuracy-efficiency dilemma. To address this challenge, we propose TALRanker, a novel framework that formalizes pointwise relevance scoring as an agentic Markov decision process. We optimize it via a two-stage training paradigm. An initial warm-up utilizes a language-preserving hybrid loss to prevent the catastrophic forgetting of native generative capacities. Subsequently, an asymmetric cost-aware reward equipped in reinforcement learning forces the policy to autonomously bypass tools for maximum efficiency when confident, while selectively retrieving external evidence to avert severe hallucination penalties when uncertain. Extensive evaluations demonstrate that TALRanker achieves state-of-the-art performance across standard and reasoning-intensive retrieval benchmarks, matching throughput with pointwise rerankers while outperforming parameter-heavy reasoning models.
Comments12 pages, 7 figures, 6 tables