arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30471cs.CLcs.AIcs.LG

CARGO:生产环境中智能体AI的上下文感知检索门控评估

CARGO: Context-Aware Retrieval-Gated Evaluation of Agentic AI in Production

Mukul Chhabra, Shail Patel, Luigi Medrano

首次发表
浏览论文内容

中文总结 AI 辅助

针对智能体系统评估中参考实例发散问题,提出CARGO框架,通过上下文接地和检索置信度门控,消除错误惩罚并提高区分性,同时暴露了宽容性带来的检测盲点。

中文摘要 AI 辅助

基于参考的LLM-as-a-judge评估假设参考答案是目标。在部署的智能体系统中,这些系统操作动态实体(支持案例、资产、账户),最接近的可用参考通常将正确的程序应用于不同的实体,因此字面判断会将不同的标识符、日期和状态视为错误或幻觉。我们将这种失败模式命名为参考实例发散(RID)。我们提出CARGO框架,该框架(i)将检索到的参考视为程序性范例,并将事实判断基于实时实例的观察上下文;(ii)为每个声明分配三种状态(支持、矛盾、无法验证),并且仅惩罚矛盾;(iii)通过检索置信度门控评估,将生产评估转化为选择性预测。我们引入了CARGO-Bench,一个基于扰动的诊断套件,其真值由构造保证,将宽容性与区分性分离。在CARGO-Bench上(246个项目,两个评判模型,7,872个判断),标准的基于参考的评判者惩罚了100%的正确实体移植答案,并且无信息量(区分指数DI约0);提供实时事实而不重新框定则没有改变任何内容。CARGO消除了这些错误惩罚(0/50),同时保持了近乎完整的矛盾召回率(50/50和49/50),将DI提高到0.58 [0.48, 0.68];一个rubric-swap对照将大部分效果归因于上下文接地维度定义。CARGO也暴露了其自身设计的一个局限性:保护实体值的宽容性抑制了对程序性腐败的检测(20%召回率)。事后修复并未缩小差距,而一项带有书面指南和裁决的LLM-as-annotator研究显示了相同的盲点。我们发布了一个预注册协议,用于将评估扩展到专家一致性、风险覆盖率和生产流量的成本。

英文摘要

Reference-based LLM-as-a-judge evaluation assumes the reference answer is the target. In deployed agentic systems that operate over dynamic entities (support cases, assets, accounts), the closest available reference typically applies the correct procedure to a different entity, so a literal judge penalizes different identifiers, dates, and statuses as errors or hallucinations. We name this failure mode reference-instance divergence (RID). We propose CARGO, a framework that (i) treats retrieved references as procedural exemplars and grounds factual judgments in the live instance's observed context, (ii) assigns each claim a three-way status (supported, contradicted, unverifiable) and penalizes only contradictions, and (iii) gates evaluation by retrieval confidence, casting production evaluation as selective prediction. We introduce CARGO-Bench, a perturbation-based diagnostic suite with ground truth by construction that separates leniency from discrimination. On CARGO-Bench (246 items, two judge models, 7,872 judgments), the standard reference-based judge penalizes 100% of correct entity-transplanted answers and is uninformative (discrimination index DI ~ 0); supplying the live facts without reframing changes nothing. CARGO eliminates these false penalties (0/50) while retaining near-complete contradiction recall (50/50 and 49/50), raising DI to 0.58 [0.48, 0.68]; a rubric-swap control attributes most of the effect to context-grounded dimension definitions. CARGO also exposes a limitation of its own design: the leniency that protects entity values suppresses detection of procedural corruptions (20% recall). A post-hoc fix does not close the gap, and an LLM-as-annotator study with written guidelines and adjudication shows the same blind spot. We release a preregistered protocol for extending the evaluation to expert agreement, risk-coverage, and cost on production traffic.

发表机构

  • Dell Technologies(戴尔科技)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑