arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03311cs.SEcs.AI

评估用于企业自动化的大语言模型权衡:来自生产级企业平台中工作流生成的经验教训

Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform

Xavier Wrenn, Radoslav Raykov, Aleksandar Angelov, Hirokuni Kitahara, Yuji Watanabe, Anca Sailer

首次发表
浏览论文内容

中文总结 AI 辅助

该研究通过在生产级企业平台评估6种大语言模型的工作流生成能力,发现分段式流水线可提升结构成功率,使小型模型达生产可行性,为云工程提供模型无关的自动化方案。

中文摘要 AI 辅助

企业合规管理需要快速适应不断变化的监管框架(如DORA、AI RMF、FedRAMP)和严格的修复服务水平协议(SLA)。传统静态编排器在混合云环境中常常失效,该环境下事件驱动评估要求自动化代码在几秒内适配运行时上下文。本文介绍了在生产级企业平台中评估六种大语言模型用于AI驱动工作流生成的经验教训,评估基准包括29个真实IT自动化场景、两种生成流水线架构,以及每个提示-模型-流水线配置下的8次独立运行(总计2784次运行)。初始流水线采用整体式工作流生成,结构成功率(JSON模式有效性和正确UI渲染)为31.5%-82.8%,多数模型在复杂JSON生成上存在困难。我们开发了重新设计的分段式流水线,将工作流构建分解为变量脚手架、基础块组装和嵌套块生成,使所有模型的结构成功率提升至74.1%-97.8%。我们分析了生产中的权衡因素,包括成本(每个工作流0.008-0.20美元)、延迟(交互使用时低于50秒)和模型选择。分段式分解使较小规模模型(如mistral-small,结构成功率95.7%,每个工作流成本0.01美元)达到生产可行性,摆脱了对昂贵前沿模型的依赖。尽管mistral-medium-2505和gpt-oss-120b取得最高结构成功率(96.1%和97.8%),但mistral-medium-2505的成本是mistral-small的19倍。我们的部署经验强调需区分结构有效性与语义正确性(用户意图的逻辑实现),并为云工程提供了一种模型无关、可扩展的自动化解决方案。

英文摘要

Enterprise compliance management requires rapid adaptation to evolving regulatory frameworks (e.g., DORA, AI RMF, FedRAMP) and tight remediation SLAs. Traditional static orchestrators often fail in hybrid cloud environments where event-driven assessments demand that automation code adapt to runtime context in seconds. This paper presents lessons learned from evaluating six large language models for AI-driven workflow generation in a production enterprise platform, benchmarked across 29 real-world IT automation scenarios, two generation pipeline architectures, and eight independent runs per prompt-model-pipeline configuration (2,784 runs total). Our initial pipeline used monolithic workflow generation, achieving 31.5-82.8% structural success rates (JSON schema validity and correct UI rendering), with most models struggling on complex JSON generation. We developed a redesigned piecewise pipeline that decomposes workflow construction into variable scaffolding, base block assembly, and nested block generation, raising structural success to 74.1-97.8% across all models. We analyze production tradeoffs including cost (USD 0.008-0.20 per workflow), latency (under 50s for interactive use), and model selection. Piecewise decomposition enables smaller models (e.g., mistral-small at 95.7% structural success and USD 0.01 per workflow) to reach production viability, removing dependency on expensive frontier models. While mistral-medium-2505 and gpt-oss-120b achieved the highest structural success (96.1% and 97.8%), mistral-medium-2505 carries a 19x cost premium versus mistral-small. Our deployment lessons highlight the need to separate structural validity from semantic correctness (logical fulfillment of user intent) and provide a solution for model-agnostic, scalable automation in cloud engineering.

发表机构

  • IBM(国际商业机器公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑