arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2604.04936cs.IRcs.AI

面向高效且经济的检索增强生成系统的网页检索感知分块(W-RAC)

Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems

  • AI Research Team, Yellow.ai(Yellow.ai AI研究团队)

机构由 AI 辅助整理,请以论文原文为准。

Uday Allu, Sonu Kedia, Tanmay Odapally, Biddwan Ahmed

更新

AI总结:

本文提出W-RAC,一种针对网页文档的高效分块框架,通过分离文本提取与语义分块规划,减少LLM token消耗,提升系统可观察性。

AI中文摘要:

检索增强生成(RAG)系统依赖有效的文档分块策略来平衡检索质量、延迟和运营成本。传统分块方法如固定大小、规则基于或完全代理分块常导致高token消耗、冗余文本生成、有限可扩展性和差的调试性,尤其在大规模网页内容摄入时。本文提出Web Retrieval-Aware Chunking(W-RAC),一种专为网页文档设计的新、经济高效的分块框架。W-RAC通过将解析的网页内容表示为结构化、ID可寻址的单元,并利用大型语言模型(LLMs)仅进行检索感知的分组决策而非文本生成,显著减少token使用,消除幻觉风险,并提高系统可观察性。实验分析和架构比较表明,W-RAC在检索性能上与传统分块方法相当或更优,同时将分块相关的LLM成本减少一个数量级。

英文摘要:

Retrieval-Augmented Generation (RAG) systems critically depend on effective document chunking strategies to balance retrieval quality, latency, and operational cost. Traditional chunking approaches, such as fixed-size, rule-based, or fully agentic chunking, often suffer from high token consumption, redundant text generation, limited scalability, and poor debuggability, especially for large-scale web content ingestion. In this paper, we propose Web Retrieval-Aware Chunking (W-RAC), a novel, cost-efficient chunking framework designed specifically for web-based documents. W-RAC decouples text extraction from semantic chunk planning by representing parsed web content as structured, ID-addressable units and leveraging large language models (LLMs) only for retrieval-aware grouping decisions rather than text generation. This significantly reduces token usage, eliminates hallucination risks, and improves system observability.Experimental analysis and architectural comparison demonstrate that W-RAC achieves comparable or better retrieval performance than traditional chunking approaches while reducing chunking-related LLM costs by an order of magnitude.

补充信息

↑