从生产流量到后训练:构建覆盖企业请求组合的自托管大语言模型(LLM)
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix
浏览论文内容
中文总结 AI 辅助
针对企业自托管LLM的GPU池碎片化问题,通过分维度训练GRPO专家并经SLERP合并,构建的单模型可整合50%平台流量,性能优于大7倍的基线模型。
中文摘要 AI 辅助
数据驻留约束迫使企业自托管大语言模型(LLM),但持续采用更新模型而不淘汰旧模型会扩大服务集群,导致有限的GPU池碎片化。我们通过生产错误分析确定的三个维度(指令遵循、函数调用、内部任务分配)的质量差距,将来自200多个内部应用的流量整合到单个模型上。质量通过针对生产流量分层的离线基准进行跟踪,并由确定性验证器或校准的LLM评判者评分。我们没有联合优化所有目标(这会引入跨域奖励干扰),而是为每个维度训练一个单独的GRPO专家,并通过两阶段SLERP合并它们。每个专家的奖励暴露了不同的失败模式:语义崩溃、过度调用、冗长篡改,每种模式都需要特定领域的修复。在非推理模式下,该方案在内部Arena上超过了总参数规模大7倍左右的基线模型,指令遵循得分69.6对65.8,函数调用得分0.85对0.83,同时提升了通用对话基准。该模型吸收了平台50%的流量,每月1.16亿个请求,且服务成本仅为基线的一小部分。
英文摘要
Data-residency constraints force enterprises to self-host LLMs, but continuous adoption of newer models without decommissioning their predecessors expands the serving fleet, fragmenting a finite GPU pool. We consolidate traffic from over 200 internal applications onto a single model by closing quality gaps identified through production error analysis along three axes: instruction following, function-calling, and internal task distribution. Quality is tracked by offline benchmarks stratified to production traffic and scored by deterministic verifiers or calibrated LLM judges. Rather than optimising all objectives jointly, which introduces cross-domain reward interference, we train a separate GRPO expert per axis and merge them via two-stage SLERP. Each expert's reward exposes a distinct failure mode, namely semantic collapse, over-calling, and verbosity hacking, each requiring a domain-specific fix. In non-reasoning mode the recipe surpasses a ${\sim}7\times$ larger by total parameters baseline on the in-house Arena with 69.6 to 65.8, instruction following with 0.85 to 0.83, and function-calling with 0.79 to 0.77, while lifting general dialogue benchmarks. The model absorbs 50% of platform traffic, 116M requests per month, at a fraction of the serving cost.
发表机构
- T-Tech
机构由 AI 辅助整理,请以论文原文为准。