发表机构
Protom Group S.p.A.(Protom集团股份公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出 AGO 质量门框架,通过四状态决策、分层评分和 beta-二项门量化回归风险,强调评判器质量需按参与度验证,点估计不足以支撑发布决策。
AI 中文摘要
采用检索增强生成(RAG)的企业面临一个反复出现的运营决策:提升、修订或阻止系统版本。证据不完整,且指标来自可能出错的 LLM 评判器。我们报告了 AGO AI 质量门(AGO),这是一个在工业 RAG 评估项目中部署的基于证据的质量门框架。AGO 整合了四个关键组件:一个四状态决策模型,将缺失数据和评判器错误视为显式结果;分层评分,结合确定性检查、本地护栏和结构化 LLM 评估;一个分层 beta-二项门,以概率方式量化回归风险;以及一个强制性的元评估协议,在 LLM 评判器影响决策之前对其进行验证。由于参与数据是专有的,我们在 RAGBench 上评估评判器层,这是一个包含 12 个数据集中 10 万条注释 RAG 轨迹的公共基准。在相同的分层测试样本上(每个评判器 N=1200),一个低成本评判器(gpt-4.1-nano)检测不合规答案的能力仅略高于随机水平(AUROC 0.603 [0.570, 0.634]),尽管其协议输出完美无缺,而 gpt-4o 达到 0.783 [0.756, 0.807]——但其各领域性能仍从 0.62 到 0.88 不等。一项涵盖回归、无变化和改进的固定种子门研究量化了不安全提升、误报成本和改进吞吐量。在回归下,决策级配置将不安全提升降至 22.2%-35.1%,而朴素门为 29.3%-41.8%。这些结果支持了设计选择:评判器质量必须按参与度衡量,且仅凭点估计不能作为发布决策。
英文摘要
Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or block a system version. The evidence is incomplete and the metrics come from fallible LLM judges. We report on AGO AI Quality Gate (AGO), an evidence-first quality-gate framework deployed in industrial RAG assessment engagements. AGO integrates four key components: a four-state decision model that treats missing data and judge errors as explicit outcomes; layered scoring combining deterministic checks, local guardrails, and structured LLM evaluation; a stratified beta-binomial gate that quantifies regression risk probabilistically; and a mandatory meta-evaluation protocol to validate the LLM judge before it influences decisions. Since engagement data is proprietary, we evaluate the judge layer on RAGBench, a public benchmark of 100k annotated RAG traces across 12 datasets. On identical stratified test samples (N=1200 per judge), a low-cost judge (gpt-4.1-nano) detects non-adherent answers barely above chance (AUROC 0.603 [0.570, 0.634]), despite producing flawless protocol output, while gpt-4o reaches 0.783 [0.756, 0.807] -- yet its per-domain performance still ranges from 0.62 to 0.88. A fixed-seed gate study spanning regression, no change, and improvement quantifies unsafe promotion, false-alarm cost, and improvement throughput. Under regression, the decision-grade profile reduces unsafe promotion to 22.2%-35.1%, against 29.3%-41.8% for a naive gate. These results support the design choices that judge quality must be measured per engagement and that point estimates alone are not a release decision.
Comments14 pages, 1 figure, 4 tables. Submitted version (pre-review). Accepted at NFMCP 2026, ECML PKDD 2026 Workshops