发表机构
Centre for Artificial Intelligence; ZHAW School of Engineering(人工智能中心; 苏黎世应用科技大学工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对RAG系统评估的不足,提出RAT统一贝叶斯框架,联合建模多维度指标,通过27种配置验证其能揭示系统差异,还分析了注释分配并扩展模型以结合人工与自动评估。
AI 中文摘要
评估检索增强生成(RAG)系统,不仅需要评估端到端的正确性,还需评估各组件的交互方式以及错误如何在流程中传播。我们提出了一种贝叶斯评估框架,该框架根据流程的信息流对检索成功、弃权(不执行)行为和答案正确性进行联合建模。该模型区分任务成功——用户是否收到正确答案(来自生成器成功)以及生成器在给定检索结果下的行为是否恰当。我们将该框架应用于三个数据集、三个检索器和三个生成器组成的27种RAG配置,结果表明,条件分解能揭示在边际指标下表现等效的系统之间存在的显著行为差异。我们进一步分析了注释分配问题,证明对于策略依从性估计,检索成功注释比任务成功注释更具信息量,并为这种不对称性提供了信息论解释。最后,我们扩展该模型以将LLM-as-a-judge注释作为校准的噪声观测值纳入,使从业者能在统一概率模型中结合有限的人工判断与成本更低的自动评估。
英文摘要
Evaluating Retrieval-Augmented Generation (RAG) systems requires assessing not only end-to-end correctness but also how individual components interact and how errors propagate through the pipeline. We introduce a Bayesian evaluation framework that jointly models retrieval success, abstention behavior, and answer correctness, factorized according to the pipeline's information flow. The model distinguishes task success. Whether the user received a correct answer (from generator success) and whether the generator behaved appropriately given the retrieval outcome. We apply the framework to 27 RAG configurations across three datasets, three retrievers, and three generators, and show that the conditional decomposition reveals substantial behavioral differences between systems that appear equivalent under marginal metrics. We further analyze the annotation allocation problem, demonstrating that retrieval-success annotations are more informative than task-success annotations for estimating policy adherence, and provide an information-theoretic explanation for this asymmetry. Finally, we extend the model to incorporate LLM-as-a-judge annotations as calibrated noisy observations, enabling practitioners to combine limited human judgments with cheaper automated assessments within a unified probabilistic model.