超越能力基准:从生产事件元数据学习LLM云服务的操作指纹
Beyond Capability Benchmarks: Learning Operational Fingerprints of LLM Cloud Services from Production Incident Metadata
浏览论文内容
中文总结 AI 辅助
该研究提出OpEmbed框架,利用生产支持案例元数据学习LLM云服务的操作指纹,在Google Cloud的大规模生产案例评估中表现优异,可用于模型上线、支持评估等场景。
中文摘要 AI 辅助
托管式LLM服务现已成为实际生产系统的一部分,但模型选择和服务规划仍严重依赖能力基准,这些基准几乎无法揭示部署后的操作行为。我们提出了操作嵌入(OpEmbed)框架,该框架无需使用案例文本,仅通过结构化且隐私保护的支持案例元数据学习LLM云服务的紧凑操作指纹。OpEmbed将模型-时间窗口聚合为八通道操作特征,并通过时间对比学习、跨视图重建和代际序数正则化学习低维表示。在Google Cloud涵盖7个LLM系列、26个月内超过33,000个生产支持案例的评估中,OpEmbed可恢复可解释的系列和版本级结构,在留一模型操作预测任务中优于非学习基线,在早期窗口数据有限时仍有用,且支持跨模型故障类型迁移。我们报告了构建和评估此工具的实践经验,该工具可用于模型上线、支持就绪性评估和操作监控。
英文摘要
Managed LLM services are now part of real production systems, but model selection and service planning still rely heavily on capability benchmarks that reveal little about operational behavior after deployment. We present Operational Embedding (OpEmbed), a framework for learning compact operational fingerprints of LLM cloud services from structured, privacy-preserving support-case metadata, without using case text. OpEmbed aggregates model--time windows into an eight-channel operational signature and learns a low-dimensional representation via temporal contrastive learning, cross-view reconstruction, and generational-ordinality regularization. Evaluated on more than 33,000 production support cases spanning seven LLM families over 26 months at Google Cloud, OpEmbed recovers interpretable family- and version-level structure, improves leave-one-model-out operational forecasting over non-learned baselines, remains useful under limited early-window data, and supports cross-model fault-type transfer. We report the practical lessons learned from building and evaluating this tool for model onboarding, support readiness assessment, and operational monitoring.
发表机构
- Google Cloud Platform, Google LLC(谷歌云平台,谷歌有限责任公司)
机构由 AI 辅助整理,请以论文原文为准。