arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2605.14488cs.AI

Deepchecks: 评估检索增强生成(RAG)

Deepchecks: Evaluating Retrieval-Augmented Generation (RAG)

  • Deepchecks, Ramat Gan, Israel(深检查,以色列拉马特甘)
  • Ben-Gurion University, Beer Sheva, Israel(本· Gurion大学,以色列贝尔谢巴)

机构由 AI 辅助整理,请以论文原文为准。

Assaf Gerner, Netta Madvil, Nadav Barak, Alex Zaikman, Jonatan Liberman, Liron Hamra, Rotem Brazilay, Shay Tsadok, Yaron Friedman, Neal Harow, Noam Bresler, Shi… 展开作者

Assaf Gerner, Netta Madvil, Nadav Barak, Alex Zaikman, Jonatan Liberman, Liron Hamra, Rotem Brazilay, Shay Tsadok, Yaron Friedman, Neal Harow, Noam Bresler, Shir Chorev, Philip Tannor, Lior Rokach

更新

AI总结:

本文提出Deepchecks框架,用于评估RAG系统,通过多维方法和根本原因分析,提升系统可靠性、相关性和用户满意度。

AI中文摘要:

大型语言模型(LLMs)结合检索增强生成(RAG)技术正在多个领域如医疗、金融和客户服务中革新应用。尽管潜力巨大,但评估RAG系统仍面临挑战,因生成输出的随机性和检索与生成组件的复杂交互。本文介绍Deepchecks,一种针对RAG应用的综合框架,通过多维方法和根本原因分析,确保与应用特定需求一致,为评估RAG系统的可靠性、相关性和用户满意度提供坚实基础。

英文摘要:

Large Language Models (LLMs) augmented with Retrieval-Augmented Generation (RAG) techniques are revolutionizing applications across multiple domains, such as healthcare, finance, and customer service. Despite their potential, evaluating RAG systems remains a complex challenge due to the stochastic nature of generated outputs and the intricate interplay between retrieval and generation components. This paper introduces Deepchecks, a comprehensive framework tailored for evaluating RAG applications. Deepchecks' evaluation framework addresses RAG applications evaluation through a multi-faceted approach, root cause analysis and production monitoring. By ensuring alignment with application-specific requirements, Deepchecks framework provides a robust foundation for assessing reliability, relevance, and user satisfaction in RAG systems.

↑