arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

释放大语言模型的潜力:面向实时、企业级部署的蓝图

Unleashing the Potential of Large Language Models: A Blueprint for Real-Time, Enterprise-Ready Deployments

Muhammad Faizan Raza, Shuo, Yang, Satish Mahadevan Srinivasan, Joanna F. DeFranco

arXiv 2608.00419首次发表:更新:

AI 中文总结

针对实时受监管场景中大语言模型的问题,提出基于模式的LLMOps架构,整合多模块并实现四项模式化贡献,优化权衡同时支持高风险领域的可审计与回滚部署。

AI 中文摘要

部署在实时、受监管场景中的大语言模型面临知识过时、灾难性遗忘、幻觉以及反馈循环薄弱等问题。我们提出一种统一的、基于模式的LLMOps架构,将实时数据摄入、持续学习、检索增强生成(RAG)以及人在回路反馈整合为单一操作流程。四项贡献对应成熟的软件设计模式:经FreshStreamBench评估的自适应摄入模式编排器(AIPO);具备稀疏时序适配器路由和感知新鲜度重放的STAR+FAR持续学习;基于服务水平目标(SLO)的自适应检索策略SAGE,可预测单查询段落预算以满足尾部延迟目标;以及带有RLHF触发机制的自动化反馈驱动收敛阶段。该架构在降低延迟-成本-精度权衡的同时,支持医疗、金融等高风险领域的可审计性与回滚功能。

英文摘要

Large language models deployed in real-time, regulated settings face knowledge staleness, catastrophic forgetting, hallucination, and weak feedback loops. We present a unified, pattern-driven LLMOps architecture integrating real-time data ingestion, continual learning, retrieval-augmented generation (RAG), and human-in-the-loop feedback into a single operational pipeline. Four contributions map to established software design patterns: an adaptive ingestion pattern orchestrator (AIPO) evaluated with FreshStreamBench; STAR+FAR continual learning with sparse temporal adapter routing and freshness-aware replay; SAGE, an SLO-aware adaptive retrieval policy predicting a per-query passage budget to meet tail-latency targets; and an automated feedback-driven convergence stage with RLHF triggers. The result reduces latency-cost-accuracy trade-offs while supporting auditability and rollback for high-risk sectors such as health care and finance.

Comments6 pages, 1 figure. Authors' accepted version of an article published in IEEE Computer. The version of record is available at the DOI below

Journal refComputer, vol. 59, no. 4, pp. 195-199, April 2026

DOI:10.1109/MC.2026.3664470

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑