搜索、检查、获取:利用布尔检索构建深度研究智能体
Search, Inspect, Fetch: Exploiting Structure-Aware Boolean Retrieval for Deep-Search Agents
浏览论文内容
中文总结 AI 辅助
本研究针对现有深度研究智能体的缺陷,提出基于BQL的SIEVE接口,在三个问答数据集上实现更高准确率且token用量显著减少,验证了BQL过滤的有效性。
中文摘要 AI 辅助
现有深度研究智能体采用搜索-访问工作流,检索并读取完整网页,未考虑网页通过标题、标题、章节和元数据暴露的可寻址结构,这导致智能体无法将检索直接限制在文档字段内,且常将不相关页面内容带入上下文。我们提出SIEVE,一种由字段化布尔检索(BQL)驱动的搜索-检查-获取接口,SIEVE会在文档字段上过滤候选对象,对准入集合排序,呈现结构丰富的结果卡片供检查,且仅获取选定章节。在三个问答数据集上,SIEVE的准确率均高于各数据集上最准确的传统搜索-访问配置,同时使用的token减少20.7%-50.6%。进一步分析显示,BQL过滤可提升所有测试的排序器性能,且准确率-上下文优势在不同检索器选择和智能体骨干模型中均持续存在。代码和数据可在此https URL获取。
英文摘要
Existing deep-search agents use a Search-Visit workflow that retrieves whole webpages without considering the structure they expose through titles, headings, sections, and metadata. This prevents agents from directly constraining retrieval to parts of a webpage and often carries irrelevant content into their context. We introduce Sieve, a search-inspect-fetch strategy driven by a Boolean Query Language (BQL): it searches webpage fields to filter candidates, uses an interchangeable ranker to order them, presents structure-rich result cards for inspection, and fetches only selected sections. Across three QA collections, Sieve is more accurate than the strongest conventional Search-Visit configuration on each collection while using 20.7-50.6% fewer tokens. Boolean filtering improves every tested ranker, and the accuracy-context advantage persists across retriever choices and agent backbones. Our implementation is included in the SkimSearchAgent library https://github.com/ielab/skim-search-agent.
发表机构
- The University of Queensland(昆士兰大学)
- CSIRO(联邦科学与工业研究组织)
机构由 AI 辅助整理,请以论文原文为准。