检索何时失效?评估用于农业咨询的RAG架构
Where Does Retrieval Fail? Evaluating RAG Architectures for Agricultural Advisory
浏览论文内容
中文总结 AI 辅助
该研究针对孟加拉语农业咨询场景,构建测试集评估RAG的检索架构与嵌入模型,发现单一综合分数无法反映检索差异,提出低资源RAG评估需按语言条件和查询类型报告性能。
中文摘要 AI 辅助
RAG系统中的检索质量通常以单一的综合分数报告,这可能掩盖不同查询类型和语言条件下的巨大差异。我们在孟加拉语农业咨询场景中研究该问题,农民查询常使用口语化表达,而官方咨询文档则采用正式科学术语。我们构建了包含1000个查询和从284份孟加拉国官方农业出版物中提取的2882个知识节点的测试集,并用其在三种受控语言条件下评估五种检索架构和六种嵌入模型。结果显示,没有单一检索方法始终最优:对于原生孟加拉语查询,BM25是最强的单一检索器(R@10=0.506),而Hybrid RRF的总体R@10最高,为0.539;但密集检索性能随查询类型变化显著,口语化农民查询的R@10为0.093,正式安全查询的R@10达0.970。在语言条件方面,BM25的R@10从孟加拉语查询的0.506降至英语查询与孟加拉语语料匹配时的0.004,而密集检索仅从0.464降至0.425。我们还发现,嵌入任务配置和段落长度可使报告的R@10变化达7倍,且与架构无关。这些结果表明,低资源RAG评估应按语言条件和查询类型报告性能,而非仅依赖综合分数。该数据集和评估脚本可在指定URL获取。
英文摘要
Retrieval quality in RAG systems is commonly reported as a single aggregate score, which can hide large differences across query types and language conditions. We study this problem in Bengali agricultural advisory, where farmer queries are often colloquial while official advisory documents use formal scientific terminology. We construct a test collection of 1,000 queries and 2,882 knowledge nodes extracted from 284 official Bangladeshi agricultural publications, and use it to evaluate five retrieval architectures and six embedding models under three controlled language conditions. The results show that no single retrieval method is consistently best. For native Bengali queries, BM25 is the strongest single retriever (R@10 = 0.506) while Hybrid RRF reaches the highest overall R@10 of 0.539. However, dense retrieval performance varies sharply by query type: R@10 is 0.093 on colloquial farmer queries and 0.970 on formal safety queries. Across language conditions, BM25 R@10 drops from 0.506 on Bengali queries to 0.004 when English queries are matched against the Bengali corpus, while dense retrieval falls only from 0.464 to 0.425. We also find that embedding task configuration and passage length can each change reported R@10 by a factor of seven, independent of architecture. These results show why low-resource RAG evaluation should report performance by language condition and query type rather than relying on aggregate scores alone. The dataset and evaluation scripts are available at https://huggingface.co/datasets/RaiyanKhaan/AgriTrust-RAG.
发表机构
- North South University(北南大学)
机构由 AI 辅助整理,请以论文原文为准。