arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LENS:基于动态原始文档的潜在证据探索的上下文内搜索

LENS: In-Context Search via Latent Evidence Exploration over Dynamic Raw Documents

Xingjun Wang, Gongsheng Li, Qi Fan, Yunlin Mao, Luyan Su, Yingda Chen

arXiv 2608.16185首次发表:更新:

AI 中文总结

针对动态原始文档的上下文内搜索问题,提出无索引框架LENS,通过迭代选择候选证据并更新置信度实现预算约束下的证据定位,在多组实验中其证据召回率和答案依据表现优于基线方法。

AI 中文摘要

大型语言模型(LLM)智能体越来越多地针对动态原始文档集合回答问题,这类文档在预处理前可能发生变化,且相关证据(片段、章节、页面或表格)依赖于查询。现有的检索增强方法通过固定分块、嵌入或持久索引预先实例化证据,虽便于查找,但成本高、易过时,且在查询已知前就已确定粒度。我们将上下文内搜索表述为在动态原始文档诱导的潜在证据空间上的预算约束证据定位问题,并提出LENS(Latent Evidence Exploration and Search,潜在证据探索与搜索),这是一种无索引框架。LENS不预先实例化证据空间,而是维持查询条件下对候选单元的置信度,通过互补的词汇、局部和探索性提议策略迭代选择候选,利用LLM相关性预言更新置信度,并在可控预算下向高后验区域收敛。证据被整合为紧凑的、基于源的感兴趣区域,并压缩为可在相关查询间复用的自组织知识簇。在500个问题的受控评估中,使用匹配的语料库快照,LENS达到62.4%的精确匹配和84.8%的证据召回率,而ReAct风格基线的精确匹配为65.2%但证据召回率为50.4%。在不同规模下,LENS在支持事实定位和答案依据方面表现最强。在150个问题的固定全维基子集(使用原始维基转储,零索引)上,LENS与ReAct在官方答案质量上几乎持平(精确匹配率分别为43.3%和42.7%),但LENS将更多答案基于检索到的证据(84.0%对70.7%)。无检索的闭卷参考凸显了模型记忆的作用。LENS在语料库变化后即可供查询,无需预处理或持久索引,且始终保持基于源的证据定位。

英文摘要

LLM agents increasingly answer questions over dynamic raw-document collections, where files may change before preprocessing, and relevant evidence (spans, sections, pages, or tables) is query-dependent. Existing retrieval-augmented approaches pre-materialize evidence via fixed chunking, embeddings, or persistent indexes: effective for lookup, yet costly, stale-prone, and committed to a granularity before the query is known. We formulate in-context search as Budgeted Evidence Localization over a latent evidence space induced by dynamic raw documents and propose LENS (Latent Evidence Exploration and Search), an index-free framework. Instead of pre-materializing the evidence space, LENS maintains a query-conditioned belief over candidate units, iteratively selecting candidates via complementary lexical, local, and exploratory proposal policies, updating the belief via an LLM relevance oracle, and narrowing toward high-posterior regions under a controllable budget. Evidence is consolidated into compact, source-grounded regions of interest and compressed into self-organizing knowledge clusters reused across related queries. On a controlled 500-question evaluation with matched corpus snapshots, LENS reaches 62.4% exact match and 84.8% evidence recall vs. 65.2% exact match but 50.4% evidence recall for a ReAct-style baseline. Across scales, LENS gives the strongest supporting-fact localization and answer grounding. On a fixed 150-question fullwiki subset over the raw Wikipedia dump with zero indexing, LENS and ReAct are nearly tied in official answer quality (43.3% vs. 42.7% EM), with LENS grounding more answers in retrieved evidence (84.0% vs. 70.7%). A no-retrieval Closed-Book reference highlights the contribution of model memory. LENS is query-ready after corpus changes, needs no preprocessing or persistent index, and preserves source-grounded evidence localization throughout.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑