发表机构
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出SemSpot语义感知服务模型,使智能体应用利用LLM推理平台现货容量,经1535个案例验证可优化成本等表现,还开发了实现该模型的相关内容。
AI 中文摘要
大语言模型(LLM)智能体正日益驱动长期运行的云推理工作负载,其中模型调用在紧急程度、冗余性、完成语义和重放成本方面存在差异。模型即服务(MaaS)平台提供多种服务模型,用于权衡成本与延迟、可用性及容量承诺,这些模型主要在请求、作业或端点范围内运行,对结合平台临时供给与智能体任务不断变化的语义的支持有限。我们提出SemSpot,一种语义感知服务模型,使智能体应用能利用LLM推理平台的现货容量。在请求层面,SemSpot允许提供商发布短期报价,涵盖成功价格、完成概率和失败通知截止时间;智能体运行时会根据当前任务状态和完成规则在这些报价中进行选择。对来自六个智能体基准的1535个案例的审计确定了四种重复出现的工作流结构,并展示该服务模型如何产生不同的成本、服务时间和回退行为。借助专门的MaaS支持,令牌级SemSpot还能在长请求内保留提供商推理状态和运行时验证的语义片段。我们开发了实现SemSpot所需的服务模型、经济边界和跨层研究议程。
英文摘要
LLM agents increasingly drive long-running cloud inference workloads in which model calls differ in urgency, redundancy, completion semantics, and replay cost. Model-as-a-Service (MaaS) platforms expose several service models for trading cost against latency, availability, and capacity commitment. These models operate primarily at request, job, or endpoint scopes and provide limited support for combining transient platform supply with the evolving semantics of an agent task. We present SemSpot, a semantics-aware service model that allows agent applications to leverage the spot capacity of LLM inference platforms. At the request level, SemSpot lets a provider publish short-lived offers over successful price, completion probability, and failure-notification deadline; the agent runtime selects among these offers using the current task state and completion rule. An audit of 1,535 cases from six agent benchmarks identifies four recurring workflow structures and shows how this service model may produce different cost, service-time, and fallback behavior. With specialized MaaS support, token-level SemSpot further preserves provider inference state and runtime-verified semantic segments inside a long request. We develop the service model, economic boundary, and the cross-layer research agenda required to realize SemSpot.