arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18123cs.AI

AutoTuneBench:LLM 服务引擎智能体自动调优的可信测量

AutoTuneBench: Trustworthy Measurement for Agent Auto-Tuning of LLM Serving Engines

Li Chen

首次发表
浏览论文内容

中文总结 AI 辅助

针对LLM服务引擎自动调优中测量不可信的问题,提出AutoTuneBench基准与协议,通过代码冻结、验证器、反作弊和预注册统计实现可信测量,并发布开放工件。

中文摘要 AI 辅助

大语言模型智能体通过“提出、测量、保留”的闭环来调优 GPU 内核和服务引擎,但该闭环背后的测量并不可信。我们基于一个为期四天、包含 619 次模型调用的试点语料库,刻画了四种失效模式:稻草人基线人为制造加速效果、绝对时间无法跨机器迁移、饱和任务使比较失效、基础设施缺陷冒充科学成果。我们提出 AutoTuneBench,一个将可信性构建为架构特性的基准测试与测量协议。该协议以代码形式冻结,并带有测试强制的来源追踪;数据库级验证器拒绝不符合协议的结果;反作弊检查在智能体的修改面之外运行;比较遵循预注册的读数;测量锚定于外部已发表结果,并基于配对种子统计,且设有 5% 的跨运行变异系数上限。诚实的测量改写了结论:我们最好的内核相对于朴素基线读数为 10.6 倍,但相对于诚实基线仅为 2.03 倍;一个配置在一台机器上实现 1.174 倍,在另一台上为 1.0049 倍;一项预注册的开/关比较在共享墙钟时间上无效(2.4840 对 2.4957 毫秒);KernelBench 一级套件接受 51% 的任务,其中相对 PyTorch eager 模式的中位加速为 1.0001 倍。该协议、双引擎语料库(vLLM 和 SGLang)及其审计追踪已作为开放工件发布。

英文摘要

Large language model agents tune GPU kernels and serving engines through a closed loop of propose, measure, and keep, but the measurements behind this loop are not trustworthy. We characterize four failure modes from a four-day pilot corpus of 619 model calls: strawman baselines manufacture speedups, absolute times do not transfer across machines, saturated tasks nullify comparisons, and infrastructure defects impersonate science. We present AutoTuneBench, a benchmark and measurement protocol that makes trust architectural. The protocol is frozen as code with test-enforced provenance; a database-level validator rejects out-of-protocol results; anti-cheat checks run outside the agent's modification surface; comparisons follow pre-registered readouts; and measurements anchor to externally published results, grounded in paired-seed statistics with a 5\% cross-run coefficient-of-variation cap. Honest measurement rewrites the headlines: our best kernel reads 10.6x against a naive baseline but 2.03x against the honest one; one configuration delivers 1.174x on one machine and 1.0049x on another; a pre-registered on/off comparison nulls at a shared wall (2.4840 vs 2.4957\,ms); and the KernelBench Level-1 suite admits 51\% of tasks with median speedup 1.0001x over PyTorch eager. The protocol, the two-engine corpus (vLLM and SGLang), and its audit trail are released as open artifacts.

补充信息

↑