AI 中文总结
本研究将 Jev 决策模型集成到边缘服务编排中,替代大型语言模型进行决策,在保持服务完成的同时,将客户端决策延迟降低 15.9%-26.5%,API 费用降低约 70%,验证了决策模型替换在低延迟场景中的有效性。
AI 中文摘要
自然语言服务请求在开始执行前可能需要语言模型进行决策,这会消耗请求的部分延迟预算。我们将 Jev 的面向决策的应用程序编程接口(API)集成到边缘服务编排中,以在保持服务完成的同时减少这一开销。该集成提取四个有界意图字段,并应用共享的验证器、准入策略和调度器,在整个请求时间线中考虑决策等待。我们使用实时 API 测量和建模执行,将 Jev 与一个短的结构化输出 DeepSeek 部署进行比较,然后在一个真实的两节点光学字符识别(OCR)服务上使用自托管 Qwen 和基于规则的参考进行对比。在三个连续的测量块中,Jev 将客户端决策延迟中位数降低了 15.9% 至 26.5%。在八个配对的 OCR 条件下,Jev 在七个条件中与 DeepSeek 的正确按时完成次数持平,在一个条件中超过。在没有缓存的情况下,两个系统都正确完成的请求的中位端到端延迟降低了 11.1% 至 25.3%;每次正确完成的 API 费用降低了 69.0% 至 70.6%。重复请求缓存基本消除了延迟差异。结果表明,在测试的服务路径中,决策模型替换可以降低响应延迟和 API 费用,并确定新鲜解释是延迟节省的主要机会。
英文摘要
Natural-language service requests can require a language-model decision before execution starts, consuming part of the request's latency budget. We integrate Jev's decision-oriented application programming interface (API) into edge service orchestration to reduce this overhead while retaining service completion. The integration extracts four to eight bounded intent fields and applies a shared validator, admission policy, and scheduler, accounting for decision waiting throughout the request timeline. We compare Jev, two self-hosted decision models, and three hosted large language models (LLMs) on 8,280 verified requests and on a live admission path with modeled execution and a real optical character recognition service. Across 33 test conditions, Jev reduces median decision latency by 22.7-64.5% relative to the fastest LLM. This latency barely moves with input size, contract width, or catalog size. On four-field contracts, Jev's API fees per correct decision are 59.7-80.9% lower at a cost of a few exact-match points, while wide contracts mark the limit of the substitution. Receiving the service catalog with each request, Jev names unseen services as accurately as known ones. On the live admission path, Jev keeps 0.91-0.95 of requests exact and on time at loads where the LLMs fall below 0.1. Since caching repeated descriptions gives the interpreters nearly the same latency, Jev's gain lies in fresh decisions. These results support decision-model substitution for latency-bound admission on bounded contracts.
Comments20 pages, 10 figures, 14 tables