发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Herschel通过按需剖析实现生产级LLM推理的持续优化,以低开销识别低效模式并指导部署优化,覆盖大量模型和加速器。
AI 中文摘要
模型即服务(Model-as-a-service)平台需要持续优化,因为复杂的服务环境会暴露出部署前未发现的低效问题。详细的常开剖析(always-on profiling)会产生大量开销,而轻量级收集又遗漏了诊断所需的信息。我们提出了Herschel,一个面向生产级大语言模型(LLM)推理的持续优化系统。我们的关键洞察是,自适应的按需剖析可以在不进行持续收集的情况下提供丰富的全栈证据。Herschel能够安全地附加到选定的运行进程并从中分离,无需修改引擎或重启,并在调查揭示证据缺失时自适应地调整覆盖范围。Herschel重构算子执行和跨进程依赖关系,以识别低效机制,并使用适用的参考修复方案提出解决方案。AI代理在保留触发工作负载和依赖关系的受控条件下实现并测试引擎和内核更改,在部署前需经过专家审查。受控测试显示,主动收集的开销对于首个令牌时间(time to first token)低于0.5%,对于每个输出令牌时间(time per output token)低于7%。有界窗口(通常为30秒)避免了常开追踪的持续成本。在六个月中,Herschel收集了约17,000个追踪,覆盖超过120个模型变体和10多种加速器类型,在23%的追踪中识别出低效模式。代表性发现指导了广泛部署的优化,包括重构同步、移除未使用的计算以及改进算子实现。
英文摘要
Model-as-a-service platforms call for continuous optimization as complex serving conditions expose inefficiencies missed before deployment. Detailed always-on profiling can incur substantial overhead, while lightweight collection omits information needed for diagnosis. We present Herschel, a continuous optimization system for production large language model (LLM) inference. Our key insight is that adaptive, on-demand profiling can provide rich full-stack evidence without continuous collection. Herschel safely attaches to and detaches from selected running processes without engine changes or restarts, and adapts coverage as investigations reveal missing evidence. Herschel reconstructs operator executions and cross-process dependencies to identify inefficiency mechanisms and suggest solutions using applicable reference fixes. AI agents implement and test engine and kernel changes under controlled conditions that preserve the triggering workload and dependencies, with expert review before deployment. Controlled tests show active-collection overhead below 0.5% for time to first token and 7% for time per output token. Bounded windows, typically 30 s, avoid the continuous cost of always-on tracing. Over six months, Herschel collected approximately 17,000 traces across over 120 model variants and more than 10 accelerator types, identifying inefficiency patterns in 23% of the traces. Representative findings guide widely deployed optimizations, including restructured synchronization, removal of unused computation, and improved operator implementations.
Comments17 pages