发表机构
PayPal(贝宝(PayPal))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对企业LLM代码助手的推理成本问题,提出T2MO数据驱动框架,通过任务分级与两级路由机制实现成本最优,可降低端到端成本并支持从静态策略到智能路由器的过渡。
AI 中文摘要
企业AI代码助手会产生大量推理开销,而单纯的token成本最小化往往无法在包含重试、升级及开发者等待时间后降低端到端成本。我们提出任务到模型优化(Task-to-Model Optimization, T2MO),一种用于优化生产代码工作流中模型选择的数据驱动方法。我们将每个开发者会话视为一项任务,可被发现、分类、分级难度、在类生产环境工具中进行基准测试,并路由至能在质量和延迟约束内完成任务的最便宜模型。该框架是一个九阶段流水线,涵盖遥测检测、分类法发现、难度分级、基准构建、候选评估、最优混合推导、预测与版本规划、分阶段路由部署以及持续治理。与以token为中心的路由规则不同,我们的目标是每项已完成任务的成本,明确计入失败升级成本。我们表明,在存在升级的情况下,该预期完成成本目标弱于token成本最小化,且推导了路由边界,即给定单元中较便宜模型要值得部署必须达到的最低通过率。决策按任务类别-难度层级的两级层次结构组织,每个单元的替代机会被汇总为按流量加权的节省瀑布,按实际美元影响对替换候选进行排名。该框架支持开发者指导、支出预测,以及从静态策略到影子模式分类器、经验证的级联,最终到智能路由器的分阶段过渡。我们以适合生产部署和未来实证研究的形式描述了该方法、优化目标、评估协议及治理循环。
英文摘要
Enterprise AI coding assistants incur substantial inference spend, and naive token-cost minimization often fails to reduce end-to-end cost once retries, escalations, and developer wait time are included. We present Task-to-Model Optimization (T2MO), a data-driven methodology for optimizing model selection in production coding workflows. We treat each developer session as a task that can be discovered, classified, graded for difficulty, benchmarked in a production-like harness, and routed to the cheapest model able to complete it within quality and latency constraints. The framework is a nine-stage pipeline spanning telemetry instrumentation, taxonomy discovery, difficulty grading, benchmark construction, candidate evaluation, optimal mix derivation, forecasting and version planning, staged routing deployment, and continuous governance. Unlike token-centric routing rules, our objective is cost per completed task, with failure escalation priced in explicitly. We show that this expected-completion-cost objective weakly dominates token-cost minimization under escalation, and we derive the routing boundary, the minimum pass rate a cheaper model must reach on a given cell to be worth deploying. Decisions are organized as a two-level hierarchy of task category difficulty tier, and per-cell displacement opportunities are aggregated into a traffic-weighted savings waterfall that ranks replacement candidates by realized dollar impact. The framework supports developer guidance, spend forecasting, and a staged transition from static policies to shadow-mode classifiers, verified cascades, and ultimately an intelligent router. We describe the methodology, optimization objective, evaluation protocol, and governance loop in a form suitable for production deployment and future empirical study.
Comments11 pages, 1 figure, 2 tables