一种面向商品级CPU硬件上低延迟LLM网络搜索的三层缓存架构
A Three-Layer Caching Architecture for Low-Latency LLM Web Search on Commodity CPU Hardware
浏览论文内容
中文总结 AI 辅助
针对LLM网络搜索中的冗余计算问题,提出三层缓存架构(会话上下文、语义查询、URL嵌入),在商品CPU硬件上实现89.3%命中率和0.1ms延迟。
中文摘要 AI 辅助
AI驱动的搜索产品,如ChatGPT搜索、Google的AI Overviews和Perplexity,提供基于实时网络结果的LLM合成答案。我们开发了OreoLook(前身为lixSearch),一个使用自动化浏览器代理和提供商路由LLM推理的开源答案引擎。其本地搜索、缓存、会话管理和嵌入栈运行在商品级CPU硬件上;答案合成由远程推理提供商执行。随着使用量的增长,会话丢失上下文,等效查询触发冗余工作,URL在会话间被重复嵌入。我们提出了一种三层缓存架构:(1)会话上下文窗口,在Redis中维护最近消息的滚动窗口,并自动溢出到Huffman压缩的磁盘归档;(2)语义查询缓存,通过嵌入向量的余弦相似度捕获改写,消除冗余的LLM调用;(3)URL嵌入缓存,跨会话去重嵌入计算。部署在单个8-vCPU Intel Cascade Lake服务器(2 GHz,32 GB RAM)上,运行30个Hypercorn工作进程,分布在三个容器化副本中,评估系统报告了89.3%的聚合Redis键空间命中率,读取延迟为0.1毫秒,内存开销仅为1.38 MB。一个后台LRU驱逐守护进程将空闲会话从Redis迁移到磁盘,并按需重新水合,使得在配置的保留策略下,对话可以在数小时或数天后恢复。
英文摘要
AI-powered search products such as ChatGPT search, Google's AI Overviews, and Perplexity provide LLM-synthesized answers grounded in live web results. We developed OreoLook (formerly lixSearch), an open-source answer engine using automated browser agents and provider-routed LLM inference. Its local search, caching, session-management, and embedding stack runs on commodity CPU hardware; answer synthesis is performed by a remote inference provider. As usage grew, sessions lost context, equivalent queries triggered redundant work, and URLs were repeatedly embedded across sessions. We present a three-layer caching architecture: (1) a Session Context Window maintaining a rolling window of recent messages in Redis with automatic overflow to Huffman-compressed disk archives; (2) a Semantic Query Cache catches rephrasings via cosine similarity on embedding vectors, eliminating redundant LLM invocations; and (3) a URL Embedding Cache that deduplicates embedding computations across sessions. Deployed on a single 8-vCPU Intel Cascade Lake server (2 GHz, 32 GB RAM) running 30 Hypercorn worker processes across three containerized replicas, the evaluated system reported an 89.3% aggregate Redis keyspace hit rate with 0.1 ms read latency and just 1.38 MB of memory overhead. A background LRU eviction daemon migrates idle sessions from Redis to disk and re-hydrates them on demand, enabling conversations that can be resumed hours or days later under the configured retention policy.