发表机构
Huawei Technologies Co., Ltd.(华为技术有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
InferOpt将LLM推理配置视为约束多目标黑盒优化,通过代理搜索、预算过滤和自适应收紧,在KV和MoE空间上分别削减缓存64.4%和路由对43.0%,优于现有基线。
AI 中文摘要
服务一个LLM意味着要设置数十个推理时旋钮,从每层的KV保留到每层的专家数量。实践中,这些旋钮通过机制特定的启发式方法设置,这些方法返回单一操作点,并且无法扩展到逐层的搜索空间。我们将推理配置重新定义为约束多目标黑盒优化问题,并构建了InferOpt,一个可复用的搜索框架,仅需变量边界、确定性资源成本和评估钩子。InferOpt在冻结的采样代理集上搜索,在任何模型调用前拒绝超出预算的候选,自适应地收紧预算,并在全规模数据集上重新验证帕累托代表。一个流水线覆盖了28维连续KV空间(Qwen2.5-7B)和26维离散MoE空间(DeepSeek-V2-Lite)。在KV上,后预填充剪枝将16K缓存削减了64.4%,TPOT削减了22.9%至48.5%,并且InferOpt搜索的逐层预算比匹配的均匀预算在Full KV参考点上分别高出7.3%和14.0%。在MoE上,InferOpt搜索的top-k调度移除了43.0%的路由token-专家对,同时保持在默认值的0.59分以内,比匹配预算的基线更接近。与随机搜索、NSGA-II和MOTPE相比,InferOpt在两个空间上都领先,在KV上取得最佳代理超体积和最低保留率,在MoE上以最低专家数保持最接近未压缩参考。
英文摘要
Serving an LLM means setting dozens of inference-time knobs, from per-layer KV retention to per-layer expert counts. Practice sets them with mechanism-specific heuristics that return a single operating point and do not scale to layer-wise search spaces. We recast inference configuration as constrained multi-objective black-box optimization and build InferOpt, a reusable search framework that requires only variable bounds, a deterministic resource cost, and an evaluation hook. InferOpt searches on a frozen sampled proxy set, rejects over-budget candidates before any model call, tightens the budget adaptively, and re-validates Pareto representatives on full-scale dataset. One pipeline covers a 28-dimensional continuous KV space (Qwen2.5-7B) and a 26-dimensional discrete MoE space (DeepSeek-V2-Lite). On KV, post-prefill pruning cuts the 16K cache by 64.4% and TPOT by 22.9--48.5%, and the searched layer-wise budget by InferOpt beats a matched uniform budget by 7.3% and 14.0% of the Full KV reference points. On MoE, a searched top-k schedule by InferOpt removes 43.0% of routed token--expert pairs while staying within 0.59 points of the default, closer than the matched-budget baselines. Against Random Search, NSGA-II, and MOTPE, InferOpt leads on both spaces, taking the best proxy hypervolume and the lowest retention on KV and staying closest to the uncompressed reference at the lowest experts on MoE.