arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越Top-K:用可解释智能体操作替代黑盒检索

Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic Operations

Sagar Tamang, Ayush Vyas, Tabarakul Hazarika

arXiv 2608.06305首次发表:更新:

AI 中文总结

该研究针对长文档检索的结构缺陷,提出无嵌入智能体检索框架READ,在政府财务报告数据集上显著优于密集检索,且性能提升源于操作界面,同时发现BM25与READ无统计差异。

AI 中文摘要

针对长文档的检索增强生成目前主流设计为:将文本分块、对分块进行嵌入,然后返回与查询最接近的Top-K个邻居。我们认为,对于财务报表、审计报告、监管申报表这类重要文档,该设计在结构上存在缺陷,且我们将该论点量化验证。在一份780页的政府财务报告中,86.8%的内容行是表格行,数千个几乎相同的数值在同一嵌入空间中竞争,且一个数值的单位通常继承自其上方中位数13行的表头——因此分块边界常会将一个数值与其所属的“十万卢比(lakh)”或“千万卢比(crore)”单位分隔,产生两个数量级的错误。我们构建了一个作为钢人(steelman)的感知表格分块器,该分块器解决了单位问题,但在我们尝试的所有分块大小下,仍有27%-30%的数值分块没有财年表头。我们提出READ(Reliable Embedding-free Agentic Document-search,可靠的无嵌入智能体文档搜索),其中智能体通过三种确定性操作读取原始文档:归一化词汇搜索、结构导航和限定范围的读取,这些操作通过模型上下文协议(Model Context Protocol)对外暴露,因此其轨迹是可复现的审计追踪,而非不透明的相似度分数。在51个经过验证的问题上,READ的回答准确率为58.8%,而密集检索的准确率仅为15.7%(p_Holm=2×10^-5);即使对READ进行调优,其准确率为35.3%,仍比密集检索高出23.5个百分点(p_Holm=0.017)。使用相同循环但搭配Top-K工具的智能体仅达到27.5%的准确率,说明性能提升源于操作界面而非迭代方式。我们还报告了证据不支持的结论:BM25与READ的性能在统计上无显著差异,因此我们的结果区分了基于嵌入的检索与无嵌入检索,而非智能体与词汇搜索。

英文摘要

Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important class of documents -- financial statements, audit reports, regulatory returns -- this design is structurally unsound, and we make the argument measurable. On a 780-page government financial report, 86.8% of content lines are table rows, thousands of near-identical figures compete in one embedding space, and a figure inherits its unit from a header a median of 13 lines above it -- so a chunk boundary routinely separates a number from whether it is in lakh or crore, an error of two orders of magnitude. A table-aware chunker built as a steelman fixes the unit problem but leaves 27-30% of numeric chunks with no fiscal-year header at every chunk size we tried. We propose READ (Reliable Embedding-free Agentic Document-search), in which an agent reads the raw document through three deterministic operations -- normalized lexical search, structural navigation, and bounded span reads -- exposed over the Model Context Protocol, so a trajectory is a replayable audit trail, not an opaque similarity score. On 51 verified questions READ answers 58.8% against dense retrieval's 15.7% (p_Holm = 2 x 10^-5) -- or 35.3% tuned, which READ still leads by 23.5 points (p_Holm = 0.017). An agent given the same loop but a top-k tool reaches only 27.5%, locating the gain in the interface rather than in iteration. We also report what the evidence does not support: BM25 is statistically indistinguishable from READ, so our result separates embedding-based from embedding-free retrieval, not agentic from lexical search.

Comments20 pages, 5 figures, 14 tables. Code, benchmark, and full result trajectories: https://github.com/twospoon/READ

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑