发表机构
ELLIS Institute Tübingen, Max Planck Institute for Intelligent Systems, Tübingen AI Center(图宾根ELLIS研究所、马克斯·普朗克智能系统研究所、图宾根人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在为开放式大语言模型推理优化建立基准测试InferenceBench,让代理部署推理服务器优化推理速度。通过多种场景和配置测试,发现代理虽能提升性能,但仍有局限,瓶颈在于提出多样配置等能力,反映了代理在开放式环境中的操作能力。
AI 中文摘要
人工智能代理越来越多地用于自动化研发任务,但现有基准测试通常在规定的工作流程或狭窄的动作空间上对它们进行评估。即使是名义上的开放式任务,通常也可以通过检索知名方法并调整几个超参数来解决,这使得不清楚强大的结果是反映真正的优化还是记忆的解决方案。我们引入了InferenceBench,在其中代理必须部署与OpenAI兼容的推理服务器并优化大语言模型推理的速度。每个代理会收到一个目标大语言模型、一个H100 GPU、一个优化场景以及两小时的时间预算。三种优化场景隔离了推理的不同瓶颈(预填充延迟、解码延迟和并发请求吞吐量),第四种场景同时平衡了这三者。在15种前沿代理配置中,代理在朴素的PyTorch基线基础上可靠地实现了提升(高达8.08倍),并且常常与默认设置的服务引擎相匹配或超越(vLLM为4.05倍),但在相同时间预算下仍低于简单的超参数搜索(高达11.53倍)。对代理轨迹的定性分析表明,尽管代理列举了许多相关的优化技术,但它们绝大多数都集中在单个推理框架上。它们只测试了少数不同的配置,并将剩余预算用于重新测量、修复或优化超参数,而不是探索实质上不同的策略。这表明瓶颈不是领域知识,而是提出多样配置、系统评估它们并提交最佳识别解决方案的能力。总体而言,InferenceBench反映了代理在开放式人工智能工程环境中的操作能力,在这种环境中,记忆的解决方案带来的改进有限。
英文摘要
AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces. Even nominally open-ended tasks can often be solved by retrieving a well-known recipe and tuning a few hyperparameters, making it unclear whether strong results reflect genuine optimization or memorized solutions. We introduce InferenceBench, where an agent must deploy an OpenAI-compatible inference server and optimize the speed of LLM inference. Each agent receives a target LLM, one H100 GPU, an optimization scenario, and a wall-clock time budget of two hours. Three optimization scenarios isolate distinct bottlenecks of inference (prefill latency, decode latency, and concurrent request throughput) and a fourth balances all three at the same time. Across 15 frontier agent configurations, agents reliably improve over a naive PyTorch baseline (up to $8.08\times$) and often match or exceed serving engines with default settings ($4.05\times$ for vLLM), but still fall below a simple hyperparameter search under the same time budget (up to $11.53\times$). Qualitative analysis of agent trajectories shows that although agents enumerate many relevant optimization techniques, they overwhelmingly converge on a single inference framework. They test only a few distinct configurations and spend the remaining budget re-measuring, repairing, or optimizing hyperparameters rather than exploring substantially different strategies. This suggests the bottleneck is not domain knowledge, but the ability to propose diverse configurations, evaluate them systematically, and submit the best identified solution. Overall, InferenceBench reflects the ability of agents to operate in an open-ended AI engineering setting, where memorized solutions lead to limited improvements.