发表机构
National Institute of Technology Srinagar(斯利那加国家理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对检索增强生成的扁平索引缺陷,提出语义压缩树(SCT)分层索引,在QASPER数据集上验证其性能,发现残差表示有效但自上而下路由表现不佳。
AI 中文摘要
检索增强生成大多依赖扁平、固定粒度的索引:文档被切割成均匀的块并通过相似度进行检索,忽略了源数据的分层结构。我们提出语义压缩树(Semantic Compression Trees, SCT),这是一种分层索引,其中每个节点仅存储其语义残差——即它相对于父节点所增加的信息,检索过程从根节点逐步向下进行,因此每个查询的成本由树的深度而非集合的大小决定。我们在QASPER数据集(50篇论文,173个问题)上进行评估,采用两种仅在基准是否提供相关文档上存在差异的协议,全程使用自助法置信区间和配对显著性检验。结果喜忧参半,我们如实报告。当提供文档时,采用零大语言模型(LLM)抽取式压缩器的SCT在答案质量上与密集检索相当(F1值为0.274对0.277,p=0.37),同时使用的上下文标记减少了30%,且构建索引无需调用LLM;残差存储优于在每个节点存储完整摘要(0.274对0.205,p<0.001)。将集合扩大50倍时,扁平检索的每个查询评分工作量增加48.9倍,而SCT仅增加6.4倍。逐步向下检索本身不受支持。当提供文档时,不使用树而检索相同残差的表现相同(p=0.27);当系统必须选择文档时,向下检索的表现明显更差(0.122对0.165,p<0.001)。路由精度揭示了原因:向下检索选择正确论文的概率为20.2%,而扁平检索为39.3%,因为该选择是从根残差做出的,而根残差是树中压缩程度最高的节点。我们得出结论:残差表示值得保留,而自上而下的路由不值得。
英文摘要
Retrieval-augmented generation relies mostly on flat, fixed-granularity indexes: documents are cut into uniform chunks and retrieved by similarity, discarding the hierarchical structure of the source. We introduce Semantic Compression Trees (SCT), a hierarchical index in which each node stores only its semantic residual -- the information it adds beyond its parent -- and retrieval proceeds by progressive descent from the root, so that per-query cost is governed by tree depth rather than collection size. We evaluate on QASPER (50 papers, 173 questions) under two protocols differing only in whether the benchmark supplies the relevant document, with bootstrap confidence intervals and paired significance tests throughout. The results are mixed and we report them as such. When the document is given, SCT with a zero-LLM extractive compressor matches dense retrieval on answer quality (0.274 vs. 0.277 F1, $p = 0.37$) using 30% fewer context tokens and no LLM calls to build the index, and residual storage beats storing full summaries at each node (0.274 vs. 0.205, $p < 0.001$). Increasing the collection fifty-fold multiplies flat retrieval's per-query scoring work by 48.9x and SCT's by 6.4x. Progressive descent itself is not supported. Retrieving the same residuals without the tree performs identically when the document is given ($p = 0.27$), and descent is substantially worse when the system must select the document (0.122 vs. 0.165, $p < 0.001$). Routing accuracy localises the cause: descent selects the correct paper 20.2% of the time against 39.3% for flat retrieval, because that choice is made from the root residual, the most compressed node in the tree. We conclude that the residual representation is worth keeping and top-down routing is not.