arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14579cs.IR

超越基准分数:RAG评估中合成查询分布与真实查询分布的分歧

Beyond Benchmark Scores: How Synthetic and Authentic Query Distributions Diverge in RAG Evaluation

Filip J. Kucia, Barbara M. Gawlik

首次发表
浏览论文内容

中文总结 AI 辅助

本研究揭示RAG评估中合成查询与真实查询在长度、来源和主题上存在显著差异,导致合成基准上的高性能配置在真实查询上性能大幅下降,建议将两者视为互补的评估维度。

中文摘要 AI 辅助

RAG系统通常使用从目标文档语料库生成的合成问题集进行常规评估。虽然这种做法提供了对整体检索能力的有用检查,但仅依赖合成基准可能会在分布偏移下产生误导,并高估部署就绪度。合成生成在语料库中均匀分布问题,形成长而详细的查询;而真实用户将大部分流量集中在少数行政和程序性主题的短查询上,同时也会询问生成器从未覆盖的事项。我们在一个大学教师信息系统上展示了这一差距,比较了通过Gemini Notebook生成的1,851个合成问题与通过学生调查收集的322个真实查询。合成查询集和真实查询集存在显著差异:真实查询平均6.8个单词,而合成查询平均15.7个单词;真实查询仅来自53个唯一来源,而合成查询来自165个来源。因此,在合成基准上看似高效的配置在真实查询上会经历显著的性能下降。重要的是,针对合成查询进行优化选择了更高延迟的混合检索器。在我们的设置中,稀疏检索组件有利于长的合成问题,但对短的真实查询无益,导致延迟成本高达我们测试的最快配置的8倍。我们建议将合成查询集和真实查询集视为查询质量谱系的两个互补极端:合成数据在理想化条件下验证最大检索能力,而真实查询则测试系统对真实用户不精确、欠指定输入的鲁棒性。

英文摘要

RAG systems are routinely evaluated using synthetic question sets generated from the target document corpus. While this practice provides a useful check on overall retrieval capability, relying exclusively on synthetic benchmarks can mislead under distribution shift and overstate deployment readiness. Synthetic generation spreads questions evenly across the corpus, formulating long, detailed queries; real users put most of their traffic on a few administrative and procedural topics in short queries, while also asking about matters the generator never covers at all. We demonstrate this gap on a university faculty information system, comparing 1,851 synthetic questions generated via Gemini Notebook against 322 authentic queries collected via a student survey. The synthetic and authentic query sets differ significantly: authentic queries average 6.8 words versus 15.7 for the synthetic ones, and draw from only 53 unique sources compared to 165. Consequently, configurations that appear highly effective on synthetic benchmarks experience a substantial performance drop on authentic queries. Importantly, optimizing on synthetic queries selected a higher-latency hybrid retriever. In our setting the sparse retrieval component benefited long synthetic questions but not short authentic ones, costing up to $8\times$ the latency of the fastest configuration we tested. We propose treating synthetic and authentic query sets as complementary extremes of the query-quality spectrum: synthetic data verifies maximum retrieval capacity under idealized conditions, while authentic queries test system robustness to the imprecise, underspecified inputs of real users.

发表机构

  • Sparrows AI sp. z o.o.(Sparrows AI有限责任公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑