arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于IBM LinuxONE的Spyre加速检索增强生成:一种用于安全、高吞吐量企业AI推理的云原生架构

Spyre-Accelerated Retrieval-Augmented Generation on IBM LinuxONE: A Cloud-Native Architecture for Secure, High-Throughput Enterprise AI Inference

Sandeep Bokkasam, Pankaj D

arXiv 2608.21393首次发表:更新:

发表机构

IBM ISDL(IBM ISDL)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出一种基于IBM LinuxONE的六子系统RAG架构,利用Spyre加速器等组件实现企业AI推理,使敏感数据不离开硬件边界,延迟低于2秒,较平台外推理降低20倍,满足受监管行业的安全与吞吐量需求。

AI 中文摘要

在企业环境中运行大语言模型一直面临着实际瓶颈:数据位于一处,AI算力位于另一处,在两者之间移动敏感记录会在延迟、安全性和监管暴露方面带来诸多难题。为LinuxONE及更广泛的IBM Z系列打造的IBM Spyre加速器PCIe推理卡改变了这一局面。本文提出了一种六子系统RAG架构,该架构完全在IBM LinuxONE上运行,使用Spyre进行生成式推理,Telum II片上加速器进行轻量级分类任务,Red Hat OpenShift进行容器编排。从查询接收到向量检索、提示词组装、LLM推理、合规性过滤再到响应交付的整个流程都保留在单个LinuxONE系统内,因此敏感数据无需离开硬件边界。我们阐述了每个子系统背后的设计选择,深入分析了Spyre编译与服务栈,解释了LinuxONE的安全执行技术如何将机密计算保障扩展到AI工作负载,并将该架构与云GPU和本地替代方案进行基准测试。初步分析显示,端到端RAG延迟低于2秒,与平台外推理相比最多可降低20倍,同时保持受监管行业实际所需的强加密和可审计性态势。

英文摘要

Running large language models inside enterprise environments has always bumped up against a practical wall: the data lives in one place, the AI horsepower sits somewhere else, and moving sensitive records between the two creates real headaches around latency, security, and regulatory exposure. IBM's Spyre accelerator PCIe inference card built for LinuxONE and the broader IBM Z family changes that equation. In this paper we lay out a six-subsystem RAG architecture that runs entirely on IBM LinuxONE, using Spyre for generative inference, the Telum II on-chip accelerator for lightweight classification tasks, and Red Hat OpenShift for container orchestration. Every piece of the pipeline from query intake through vector retrieval, prompt assembly, LLM inference, compliance filtering, and response delivery stays within a single LinuxONE system, so sensitive data never has to leave the hardware perimeter. We walk through the design choices behind each subsystem, dig into the Spyre compilation and serving stack, explain how LinuxONE's Secure Execution technology extends confidential-computing guarantees to AI workloads, and benchmark the architecture against cloud-GPU and on-premises alternatives. Early analysis points to end-to-end RAG latencies under two seconds and up to a 20x reduction compared to off-platform inference, all while keeping the strong encryption and auditability posture that regulated industries actually need.

Comments8 pages, 1 figure, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑