arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FDE-Bench:评估大语言模型智能体在部署环境配置中的表现

FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration

Weihang Ding, Junfei Zhan, Yueting Li, Qirong Guo

arXiv 2609.27571首次发表:更新:

发表机构

University of California, Berkeley; Imperial College London; The Hong Kong University of Science and Technology (Guangzhou)(加州大学伯克利分校; 帝国理工学院; 香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FDE-Bench通过136个部署配置任务评估大语言模型智能体,采用四层门控检查和发布门机制,发现七个模型解决52.9%-75.0%任务,就绪状态为主要失败环节,并验证了人工指导可显著提升性能。

AI 中文摘要

部署要求智能体将应用程序代码转化为一个运行中的系统,其服务能够相互连接、达到就绪状态并保持可观测性。FDE-Bench通过136个部署配置任务来评估这一能力,任务涵盖Docker镜像、多服务Compose堆栈以及Kubernetes,并采用绿地(greenfield)和诊断修复(diagnose-and-repair)两种模式。智能体提交声明式工件,这些工件在纯净环境中被收集、重建并重新部署。四个门控二进制检查层分别衡量构建、就绪、行为及对部署规范的符合性,使用程序化检查而非LLM评判。一个四臂发布门要求存在可解析的参考解决方案,并拒绝由无操作、规范转录或通用桩提交所解决的任务。发布的检查注释揭示了2,145项检查与其规范之间的联系,包括七个已记录的缺口。三种额外的对抗策略测试评分信号中的捷径;这些策略未能解决其所覆盖的135个任务中的任何一个,而一个空洞的健康探针通过了就绪检查,暴露出对下游检查的需求。在136个任务的评估网格中,来自四个提供商的七个语言模型使用相同的四工具脚手架,解决了52.9%至75.0%的任务。三个零智能基线未解决任何任务,平均部署得分最高为0.44。就绪状态是最大的失败阶段,占313个未解决情节中的110个。修复任务组的平均解决率比不相交的绿地组高出30.7个百分点,且每个模型均呈现正差距;十个任务对所有七个模型均构成挑战。在一项25个任务的案例研究中,一位指导Claude-Sonnet-5的执业工程师解决了92%的任务,而自主基线为72%。FDE-Bench将部署成功与失败关联到可检查和可重放的工件上。

英文摘要

Deployment requires an agent to turn application code into a running system whose services connect, become ready, and remain observable. FDE-Bench evaluates this capability with 136 deployment-configuration tasks spanning Docker images, multi-service Compose stacks, and Kubernetes, in greenfield and diagnose-and-repair modes. Agents submit declarative artifacts that are collected, rebuilt, and redeployed in a pristine environment. Four gated binary check layers measure build, readiness, behavior, and conformance to the deployment specification, using programmatic checks without an LLM judge. A four-arm release gate requires a resolving reference solution and rejects tasks solved by do-nothing, specification-transcription, or generic-stub submissions. The released check annotations expose the link between 2,145 checks and their specifications, including seven documented gaps. Three additional adversarial strategies test shortcuts in the grading signals; none resolves any of the 135 tasks they cover, while a vacuous health probe passes readiness and exposes the need for downstream checks. On the 136-task evaluation grid, seven language models from four providers use the same four-tool scaffold and resolve 52.9-75.0 percent of tasks. The three zero-intelligence floors resolve none and reach a mean Deployment Score of at most 0.44. Readiness is the largest failure stage, accounting for 110 of 313 unresolved episodes. Mean resolution rate is 30.7 percentage points higher on the repair task group than on the disjoint greenfield group, with a positive gap for every model; ten tasks resist all seven. In a 25-task case study, one practicing engineer directing Claude-Sonnet-5 resolves 92 percent against 72 percent for the autonomous baseline. FDE-Bench links deployment success and failure to artifacts that can be inspected and replayed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑