arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.21841cs.AIstat.ML

EnterpriseVal:量化生成式AI在企业中的效能、可靠性与价值

EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise

Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid

首次发表
浏览论文内容

中文总结 AI 辅助

针对企业GenAI部署缺乏可测量业务效果的问题,提出用例级评估系统EnterpriseVal,通过形式化规范、指标目录、校准评分和双层门控,将指标映射至扩展决策,并在银行试点中验证其有效性。

中文摘要 AI 辅助

前沿语言模型现在能够生成专业交付物,专家评分者认为这些交付物在相当比例的经济价值任务上与人类工作相当,然而大多数企业生成式AI(GenAI)项目未能显示出可衡量的业务效果,且很大一部分智能体项目预计将被取消。我们认为这在很大程度上是一个测量问题:公共基准回答的是“模型能做什么?”,而部署决策需要回答“这个工作流是否适合、可靠、安全且值得扩展——在这里,在我们的数据上,在我们的控制下?”。我们提出了EnterpriseVal,一个用例级别的评估系统,以弥合这一差距。它包含:(i)对用例和被测的冻结的社会技术配置(包括模型、提示、检索、工具、护栏和人工监督)的正式规范,并设有自主性级别和后果等级,共同决定所需的评估强度;(ii)一个涵盖保真度、效用、效率、可靠性、保障和监督的指标目录;(iii)一个评分协议,通过预测驱动的推断,将盲法专家判断与校准的LLM-as-judge评分相结合,实现规模化;(iv)一个双层阈值门控,以可执行算法形式表述,将带有置信区间的指标向量映射到REJECT/CONDITIONAL/SCALE决策;(v)一个价值与风险模型,其中评审者捕获率是一个测量参数。我们报告了在全球银行三个工作流中的试点。在信用备忘录起草中,最佳模型的人工评分引用精确度达到88%,幻觉率为1.6%,而门控阈值分别为70%和5%;在流程转换中,分析师精炼工作量从每份文档估计的27.4小时降至2.9小时。我们区分了已确立的结果、记录的试点证据、所提出的系统和开放假设,并指定了全面验证所需的实验。

英文摘要

Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. We argue that this is substantially a measurement problem: public benchmarks answer "what can the model do?", whereas a deployment decision requires "is this workflow fit, reliable, safe and worth scaling - here, on our data, under our controls?". We present EnterpriseVal, a use-case-level evaluation system that closes this gap. It comprises (i) a formal specification of the use case and of the frozen socio-technical configuration under test, model, prompts, retrieval, tools, guardrails and human oversight, with an autonomy level and consequence tier that jointly set the required evaluation intensity; (ii) a metric catalogue spanning fidelity, utility, efficiency, reliability, assurance and oversight; (iii) a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference; (iv) a two-tier threshold gate, stated as an executable algorithm, that maps metric vectors with confidence bounds to REJECT/CONDITIONAL/SCALE decisions; and (v) a value-and-risk model in which the reviewer catch rate is a measured parameter. We report a pilot across three workflows in a global bank. In credit-memo drafting, human-graded citation precision reached 88% and hallucination rate 1.6% for the best model against gates of 70% and 5%; in procedure transformation, analyst refinement effort fell from an estimated 27.4 to 2.9 hours per document. We separate established results, documented pilot evidence, the proposed system and open hypotheses, and specify the experiments required for full validation

发表机构

  • Citigrounp, Inc(花旗集团)
  • Ernst & Young LLP(安永会计师事务所)
  • NVIDIA Corporation(英伟达公司)

机构由 AI 辅助整理,请以论文原文为准。

↑