arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.26777cs.AIcs.SE

SWE-Serve:面向生产推理服务的智能体工程基准测试

SWE-Serve: Benchmarking Agentic Engineering For Production Inference Serving

  • NVIDIA(英伟达)
  • University of California, Berkeley(加州大学伯克利分校)

机构由 AI 辅助整理,请以论文原文为准。

Jennifer Williams, Dave Farris, Jeff Farris, Jiantao Jiao

中文总结 AI 辅助

SWE-Serve是面向生产推理服务的智能体基准,含53个SGLang任务,揭示本地完成与生产正确性差距,最佳配置平均pass@1为75%。

中文摘要 AI 辅助

我们推出了SWE-Serve,一个用于在生产推理工程任务上评估智能体的基准测试。实现一个推理功能可能需要协调服务栈中的多项更改,包括模型支持、运行时执行和公共API。现有基准测试对生产推理工程的覆盖有限:仓库级软件工程基准测试不针对推理,而通用终端智能体基准测试仅包含少量推理任务。同时,专门的推理基准测试主要侧重于孤立的核生成或性能优化,而非仓库规模的生产功能实现。SWE-Serve提供了53个基于仓库的任务,这些任务源自SGLang近期生产变更,涵盖六个推理工程类别。每个任务在CPU或单个GPU(H100)上执行,并通过隐藏的功能测试和回归测试进行评估,包括适用的端到端(E2E)服务测试和校准的性能门槛。可执行的无操作(no-op)和预言机(oracle)控制、对抗性验证者审查以及闭卷执行支持任务有效性和评估完整性。在11个模型和31种模型努力配置中,性能最佳的配置达到了75%的平均pass@1。SWE-Serve揭示了在本地完成任务与实现生产正确性之间的巨大差距。在19个具有端到端覆盖的任务中,模型服务E2E测试拒绝了大约三分之一的通过所有其他测试的补丁(在验证者下为45.9%,而在评分中排除E2E测试时为69.4%),每个模型的最佳性能配置的通过率有所提高。通过使生产正确性差距可直接测量,SWE-Serve使该领域能够跟踪未来智能体是否超越本地完成任务,实现生产正确性。

英文摘要

We introduce SWE-Serve, a benchmark for evaluating agents on production inference engineering tasks. Implementing an inference feature can require coordinating multiple changes across the serving stack, including model support, runtime execution, and public APIs. Existing benchmarks provide limited coverage of production inference engineering: repository-level software engineering benchmarks do not target inference, while general terminal-agent benchmarks include only a few inference tasks. Dedicated inference benchmarks, meanwhile, focus primarily on isolated kernel generation or performance optimization rather than repository-scale production feature implementation. SWE-Serve provides 53 repository-grounded tasks derived from recent production changes to SGLang, spanning six inference engineering families. Each task executes on either CPU or a single GPU (H100) and is evaluated with hidden functional and regression tests, including, where applicable, end-to-end (E2E) serving tests and calibrated performance gates. Executable no-op and oracle controls, adversarial verifier review, and closed-book execution support task validity and evaluation integrity. Across 11 models and 31 model-effort configurations, the best-performing configuration achieves 75% mean pass@1. SWE-Serve exposes a substantial gap between completing tasks locally and achieving production correctness. On 19 tasks with end-to-end coverage, model-serving E2E tests reject roughly one-third of patches that pass every other test (45.9% under the verifier versus 69.4% with E2E tests excluded from scoring), with pass rate increasing for each model's best-performing configuration. By making the production correctness gap directly measurable, SWE-Serve enables the field to track whether future agents move beyond completing tasks locally to achieving production correctness.

↑