发表机构
University of Maryland, College Park(马里兰大学学院公园分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一个基于智能搜索的系统,利用知识图谱和大型语言模型,从自然语言查询中返回匹配的地球科学数据集和工具,显著提升检索性能。
AI 中文摘要
NASA及其数据中心拥有数千个地球科学数据集和工具,如Worldview、Giovanni、科学发现引擎和Harmony。即使是领域专家也很难找到合适的数据集。我们提出了一个智能搜索系统,作为面向地球科学社区的公共服务部署,该系统接受自然语言研究查询并返回匹配的数据集和工具。我们证明,在大语言模型时代,知识图谱的潜在价值可以通过智能搜索得到显著放大。从NASA地球观测知识图谱中,我们推导出NASA-EO-Bench,一个包含47k个查询-数据集对(21k个基于任务的查询)的开放基准。在NASA-EO-Bench上微调的神经评分器优于余弦和BM25基线。通过分数融合将其与BM25结合,Recall@10和MRR均提高了5倍以上。在此监督流水线之上,我们添加了一个零样本智能重排序阶段,无需额外训练即可在分层N=200子集上将MRR提高28%,表明LLM推理与监督检索互补。
英文摘要
NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony. Finding the right one is hard even for domain experts. We present an agentic search framework for geoscience data discovery that takes a natural-language research query and returns matching datasets and tools. We demonstrate that, in the era of large language models, the latent value of knowledge graphs (KGs) can be substantially amplified through agentic search. From the NASA Earth Observation Knowledge Graph (NASA EO-KG) we derive NASA-EO-Bench, an open benchmark of 47k query-dataset pairs (21k task-based queries). A neural scorer fine-tuned on NASA-EO-Bench beats cosine and BM25 baselines. Further combining it with BM25 via score fusion raises both Recall@10 (R@10) and MRR to over 5x the unadapted cosine baseline. On top of this supervised pipeline, a zero-shot reranking stage lifts MRR by 16%, significant under a paired bootstrap, with no additional training, and autonomous web and arXiv tool use adds a further gain, showing that LLM reasoning is complementary to supervised retrieval.
CommentsAccepted at CIKM 2026 (full research paper). v2: camera-ready version; LLM rerank model sweep extended from N=200 to N=600 test queries with paired-bootstrap significance tests; adds a fine-tuned cross-encoder (bge-reranker-v2-m3) as a supervised reranking baseline