arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22634cs.CLcs.AI

GeoRisk-RAG:一种通过选择性回答提升RAG可靠性的层级感知风险框架

GeoRisk-RAG: A Hierarchy-Aware Risk Framework for Improving RAG Reliability through Selective Answering

  • Virginia Tech(弗吉尼亚理工大学)
  • Florida Polytechnic University(佛罗里达理工大学)
  • Kuwait University(科威特大学)

机构由 AI 辅助整理,请以论文原文为准。

Meenu Ravi, Shailik Sarkar, Lulwah AlKulaib, Yordanos Tessema, Chang-Tien Lu

AI总结:

GeoRisk-RAG是一种层级感知框架,通过选择性回答结合有向无环图距离估计地理适用性,在野火QA数据集上将位置相关问题错误置信率降至0.009,提升RAG可靠性,助力地理空间领域安全决策。

AI中文摘要:

当前提升大语言模型(LLM)生成答案可靠性的研究主要利用检索增强生成(RAG)、知识图谱增强和强化学习。这些方法虽擅长通过语义相似度和忠实度提升与衡量可靠性,但往往难以区分语义相似度与地理有效性,这在自然灾害管理领域尤为关键,因为地理粒度(即镇、市、州)对决策至关重要,某一行政区有效的响应可能无法迁移至另一行政区,此类领域中,自信的错误答案比弃权(不执行)带来更大风险。本文提出GeoRisk-RAG,一种新型层级感知框架,通过选择性回答解决这一地理有效性差距,该框架在响应生成前,使用基于有向无环图(DAG)的距离进行上下文检索,明确估计地理适用性。在新的保留野火相关问答(QA)数据集上的实验显示,GeoRisk-RAG显著降低了位置相关问题的错误置信率,将其降至0.009,而标准语义相似度与重排序基线的错误置信率约为0.090,同时始终实现更高的人类偏好对齐度。本研究通过整合地理有效性与选择性回答行为,为地理空间领域更安全的决策提供了对端到端RAG管道更全面的评估。

英文摘要:

Current work on improving reliability in large language model (LLM)- generated answers has primarily leveraged Retrieval-Augmented Generation (RAG), knowledge-graph augmentation, and reinforcement learning. While these methods are adept at enhancing and measuring reliability through semantic similarity and faithfulness, they often struggle to distinguish semantic similarity from geographic validity. This is especially critical in natural hazard management domains where geographic granularity (i.e., town vs. city vs. state) is significant for decision-making, as responses valid in one municipality may not transfer to another. In such domains, a confidently wrong answer carries greater risk than abstaining. We present GeoRisk-RAG, a novel hierarchy-aware framework that addresses this geographic-validity gap through selective answering. This framework explicitly estimates geographic applicability using a Directed Acyclic Graph (DAG)-based distance for context retrieval before response generation. Experiments on a novel held-out wildfire-related question-answering (QA) dataset show that GeoRisk-RAG significantly reduces false confidence rates for location-dependent questions, lowering the rate to 0.009 compared with ~0.090 for standard semantic similarity and reranking baselines, while consistently achieving higher human preference alignment. This work provides a more comprehensive assessment of end-to-end RAG pipelines by integrating geographic validity and selective-answering behavior for safer decision-making in geospatial domains.

补充信息

↑