FinCacheServe:面向可变企业文档的高性价比RAG服务的依赖一致答案复用
FinCacheServe: Dependency-Consistent Answer Reuse for Cost-Efficient RAG Serving over Mutable Enterprise Documents
浏览论文内容
中文总结 AI 辅助
FinCacheServe为可变企业文档的RAG服务设计依赖一致的答案复用机制,可跳过大量LLM调用,相比现有方法能降低能耗,提升服务成本效率。
中文摘要 AI 辅助
针对可变企业文档的检索增强生成服务会重复执行语义等价的分析请求,答案复用可移除受GPU限制的生成工作,但响应缓存需在文件、证据块和工具输出变更时保持依赖一致性。FinCacheServe将每个生成的答案视为由企业意图索引、并受文档版本、证据指纹、工具指纹、模型身份和解码配置保护的服务对象。基于vLLM实现对SEC衍生的金融文档工作负载和Qwen2.5模型进行评估,在2230个请求的托管7B跟踪中,FinCacheServe跳过了53.27%的LLM调用,未观察到依赖陈旧的输出;在三个托管32B算子套件种子的544个请求中,其跳过了53.31%,而版本化语义缓存和 grounded 式复用分别为38.97%和22.43%;容量、后端和SLO重放显示其具备预言机约束的缓存管理、10万条事务元数据行为,且每依赖新鲜2秒SLO成功的估计Wh比版本化语义缓存低44.30%。
英文摘要
Retrieval-augmented generation services over mutable enterprise documents repeatedly execute semantically equivalent analysis requests. Answer reuse can remove GPU-bound generation work, yet response caches require dependency consistency when filings, evidence chunks, and tool outputs change. FinCacheServe treats each generated answer as a serving object indexed by enterprise intent and guarded by document versions, evidence fingerprints, tool fingerprints, model identity, and decoding configuration. A vLLM implementation evaluates SEC-derived financial-document workloads with Qwen2.5 models. On a 2,230-request hosted 7B trace, FinCacheServe skips 53.27% of LLM calls with zero observed dependency-stale outputs. Across three hosted 32B operator-suite seeds, it skips 53.31% of 544 requests, compared with 38.97% for versioned semantic caching and 22.43% for grounded-style reuse. Capacity, backend, and SLO replays show oracle-bounded cache management, 100k-entry transactional metadata behavior, and 44.30% lower estimated Wh per dependency-fresh 2s-SLO success than versioned semantic caching.
发表机构
- The Chinese University of Hong Kong(香港中文大学)
- Beijing Institute of Technology(北京理工大学)
机构由 AI 辅助整理,请以论文原文为准。