arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12188cs.AIcs.DBcs.IR

成本治理的检索增强生成模型:多租户大语言模型系统中跨检索与生成的统一租户成本归因

Cost-Governed RAG: Unified Per-Tenant Cost Attribution Across Retrieval and Generation in Multi-Tenant LLM Systems

Navnit Shukla

首次发表
浏览论文内容

中文总结 AI 辅助

研究多租户大语言模型系统中成本治理问题,提出成本治理的RAG架构,集成TurboVec与治理网关实现统一可观测堆栈,能按租户联合归因成本,准确率高且降低成本,还形式化三层成本模型。

中文摘要 AI 辅助

企业检索增强生成(RAG)部署面临关键治理差距:大语言模型生成成本按令牌计量,而检索层(向量存储、相似度计算和嵌入API调用)仍是无归属的共享成本,导致租户间存在无形的交叉补贴。我们提出成本治理的RAG,它将无码本向量索引(TurboVec)与多租户大语言模型治理网关集成,创建统一可观测堆栈,使嵌入、检索和生成成本可按租户联合归因。该架构利用TurboVec的确定性闭式内存公式实现近乎精确的按租户检索成本计算。在云数据平台治理边界内的Snowpark容器服务上部署,系统在100个模拟租户(1000万个向量,对数正态大小分布)中实现了99.96%的端到端成本归因准确率,遥测开销低于查询延迟的0.04%。与托管向量数据库服务相比,该架构将检索基础设施成本降低了3.1 - 9.0倍。我们形式化了一个三层成本模型,并证明无码本量化实现了确定性的按租户成本归因,同时消除了训练量化器中存在的共享码本泄漏面。

英文摘要

Enterprise Retrieval-Augmented Generation (RAG) deployments face a critical governance gap: while LLM generation cost is metered per token, the retrieval layer - vector memory, similarity compute, and embedding API calls - remains an unattributed shared cost, enabling invisible cross-subsidization among tenants. We present Cost-Governed RAG, an architecture that integrates a codebook-oblivious vector index (TurboVec) with a multi-tenant LLM governance gateway, creating a unified observability stack where embedding, retrieval, and generation costs are jointly attributable per tenant. The architecture exploits TurboVec's deterministic, closed-form memory formula to enable near-exact per-tenant retrieval cost calculation - a property unavailable in graph-based indexes with non-linear memory overhead. Deployed on Snowpark Container Services within a cloud data platform's governance boundary, the system achieves 99.96% end-to-end cost attribution accuracy across 100 simulated tenants (10M vectors, log-normal size distribution) with telemetry overhead below 0.04% of query latency. The architecture reduces retrieval infrastructure cost by 3.1-9.0x compared to managed vector database services under the pricing assumptions detailed in Section IV. We formalize a three-layer cost model and demonstrate that codebook-oblivious quantization enables deterministic per-tenant cost attribution while also removing the shared-codebook leakage surface present in trained quantizers - the latter observation being exploratory and subject to the limitations described in Section VII.

↑