发表机构
University of Göttingen; Bielefeld University of Applied Sciences and Arts (HSBI)(哥廷根大学; 比勒费尔德应用科学与艺术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CAFE是开源平台,用于复合人工智能系统评估。它将管道组件设为因子构建析因设计,运行配置并用模型评判器和人工评分。能归因质量差异,报告多种结果,还能解释组件影响,在问答管道验证有效,已发布为Python包和Web应用。
AI 中文摘要
我们介绍了CAFE(复合人工智能因子评估),这是一个开源平台,将实验设计引入复合人工智能系统(CAIS)的评估。此类系统有许多可互换的选择,如检索器、模型或提示等,从业者很少知道哪些对答案质量影响最大。借助CAFE,从业者将管道的每个可互换组件注册为一个因子,以在所选因子上构建析因设计,运行生成的配置,并使用可配置的语言模型评判器和人工评分者根据共享的评分标准对答案进行评分。通过这些评分,它使用混合效应模型将答案质量差异归因于组件及其相互作用,并报告效应大小、显著性、最佳配置、成本和延迟权衡以及评判器与人工的可靠性。现有工具大多要么单独搜索良好配置,要么单独对输出进行评分,而CAFE还能解释哪个组件推动了质量以及观察到的差异是否显著。我们在HotpotQA基准数据集上的检索增强问答(QA)管道上验证了CAFE,它恢复了植入的因子效应,并在排列空值下保持校准。CAFE作为Python包和Web应用程序发布。
英文摘要
We introduce CAFE (Compound-AI Factorial Evaluation), an open-source platform that brings design of experiments to the evaluation of compound AI systems (CAIS). Such systems expose many interchangeable choices - e.g. which retriever, model, or prompt - and practitioners rarely know which of them most affects answer quality. With CAFE, a practitioner registers each swappable component of a pipeline as a factor to build a factorial design over the chosen factors, run the resulting configurations, and score the answers on a shared rubric using a configurable LLM judge together with human raters. From these ratings it attributes answer-quality variance to the components and their interactions with mixed-effects models and reports effect sizes, significance, the best configuration, cost and latency trade-offs, and judge-human reliability. Whereas existing tools mostly either search for a good configuration or score outputs in isolation, CAFE also explains which component drives quality and whether an observed difference is significant. We validate CAFE on a retrieval-augmented question-answering (QA) pipeline over the HotpotQA benchmark dataset, where it recovers planted factor effects and stays calibrated under a permutation null. CAFE is released as a Python package and as a Web application.