语义交付网络:为LLM智能体重构Web检索基础设施
Semantics Delivery Network: Rethinking Web Retrieval Infrastructure for LLM Agents
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对LLM智能体检索效率低的问题,提出语义交付网络SemDN,以块粒度索引和缓存Web内容,实现跨智能体共享,初步实验显示其提升答案质量并减少冗余。
AI中文摘要:
大型语言模型(LLMs)在回答需要专有信息或最新实时网页内容的问题时,越来越依赖外部来源,这既通过传统的单次检索增强生成(RAG),也通过多轮智能体RAG。然而,当今的Web基础设施仍是为人类客户端构建的。给定一个查询,当前的搜索服务返回按通用相关性排序的URL和摘要列表;内容交付网络(CDNs)缓存以URL寻址的对象(文本、图像、视频等),而不知道智能体需要哪个段落。相比之下,LLMs消费简短、语义连贯的段落(以下称为“块”),这些块是根据下游任务效用而非仅相似性选择的,并且可能在推理轮次中进行有状态检索。不协调的智能体还会重复搜索、数据获取和语义处理,复制了本可以共享的工作。我们认为,语义块检索应成为一流的网络交付抽象。我们提出语义交付网络(SemDN):一种源授权的、分层的边缘基础设施,以块粒度索引、搜索和智能缓存Web内容。SemDN代表参与网站为智能体提供服务,跨智能体分摊数据获取和处理,并支持租户特定的检索策略。由于与URL缓存不同,语义检索不提供明确的未命中信号,SemDN必须估计其注册语料库何时可能不完整或过时,并触发有范围的发现或刷新。它引发了关于可共享检索状态、分层缓存、覆盖风险和部署的开放问题。我们的初步探索揭示了页面处理内容与消费块之间存在巨大差距,显著的局部任务重用,以及块交付带来的每个上下文令牌更高的答案质量。
英文摘要:
Large language models (LLMs) increasingly rely on external sources when answering questions that require proprietary information or up-to-date live web content, through both traditional single-shot retrieval-augmented generation (RAG) and multi-turn agentic RAG. Yet today's web infrastructure is still built for human clients. Given a query, current search services return a list of URLs and snippets ranked for generic relevance; content delivery networks (CDNs) cache URL-addressed objects (texts, images, videos, etc.) without knowing which passage an agent needs. LLMs, in contrast, consume short, semantically coherent passages, hereafter "chunks", selected for downstream task utility rather than similarity alone, and may retrieve statefully across reasoning turns. Uncoordinated agents also repeat search, data acquisition, and semantic processing, duplicating work that could be shared. We argue that semantic chunk retrieval should become a first-class network-delivery abstraction. We propose Semantics Delivery Network (SemDN): an origin-authorized, hierarchical edge substrate that indexes, searches, and smart-caches web content at chunk granularity. SemDN serves agents on behalf of participating websites, amortizes data acquisition and processing across agents, and supports tenant-specific retrieval policies. Because, unlike URL caching, semantic retrieval provides no explicit miss signal, SemDN must estimate when its enrolled corpus may be incomplete or stale and trigger scoped discovery or refresh. It raises open questions about shareable retrieval state, hierarchical caching, coverage risk, and deployment. Our preliminary probes reveal a large gap between page content processed and chunks consumed, substantial task-local reuse, and higher answer quality per context token from chunk delivery.