发表机构
Old Dominion University(奥多明尼昂大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Guardian Crawler是面向噪声网络智能的检索优先型测试平台,结合BM25与风险感知重排序等技术,在合成语料库实验中取得较高检索分数,可用于知识发现与证据摘要生成。
AI 中文摘要
从噪声网络数据中检索相关证据颇具挑战性,尤其是在包含不完整报告、异构语言及无关内容的敏感领域。本文提出了Guardian Crawler,这是一个可复现的检索优先型测试平台,用于在类合成网络语料库上开展知识发现与基于证据的摘要生成的受控实验。其架构将BM25检索与风险感知、嵌入增强及混合重排序相结合,随后进行带显式文献引用的受限检索增强生成。在包含900篇文档和10个查询的合成语料库上开展的实验显示,在基于风险的重排序下,该方法的描述性检索分数最高,P@10=1.00、NDCG@10=0.94,而BM25的对应值为0.94和0.81;最优的混合配置及BM25+Semantic配置的NDCG@10值分别达0.94和0.88。所有41个可评估的生成摘要要点均通过了词汇覆盖阈值;自动化大语言模型评判器将其中36个归类为有支持依据,1个为部分支持,4个为无支持依据。这些结果表明Guardian Crawler作为受控测试平台具备可行性,但未确立其统计优越性、经人工验证的忠实性或对实时网络调查环境的可迁移性。
英文摘要
Retrieving relevant evidence from noisy web data is challenging, particularly in sensitive domains containing incomplete reports, heterogeneous language, and irrelevant content. We present Guardian Crawler, a reproducible retrieval-first testbed for controlled experiments on knowledge discovery and evidence-grounded summarization over synthetic web-like corpora. The architecture combines BM25 retrieval with risk-aware, embedding-augmented, and hybrid reranking, followed by constrained retrieval-augmented generation with explicit document citations. Experiments on a synthetic 900-document corpus and 10 queries produced the highest descriptive retrieval scores under risk-based reranking, with P@10 = 1.00 and NDCG@10 = 0.94, compared with 0.94 and 0.81 for BM25. The best hybrid and BM25+Semantic configurations reached NDCG@10 values of 0.94 and 0.88, respectively. All 41 evaluable generated bullets passed the lexical coverage threshold; an automated LLM judge classified 36 as supported, one as partially supported, and four as unsupported. These results demonstrate the feasibility of Guardian Crawler as a controlled testbed but do not establish statistical superiority, human-validated faithfulness, or transfer to live-web investigative environments.
Comments8 pages, 2 figures. Accepted as a Short Paper at KDIR 2026 (International Conference on Knowledge Discovery and Information Retrieval)