AI 中文总结
针对长文档问答,GraphRAG虽有改进但自动构建图有缺陷。为此提出PAGE-RAG框架,将图结构视为语义骨架,引入任务自适应检索路由策略及知识边界控制,实验表明其提升了检索效率、知识可靠性与答案质量。
AI 中文摘要
GraphRAG通过引入超越传统检索的结构化表示来改进长文档问答。然而,自动构建的图本质上是源文档的不完整投影,将它们视为独立知识源可能导致不可靠的检索和生成。我们提出了PAGE-RAG,一种用于可靠长文档问答的投影感知自适应图检索框架。PAGE-RAG将图结构视为组织和导航文档知识的语义骨架,而非取代原始知识源。基于此,引入任务自适应检索路由策略,根据查询需求动态选择合适检索行为。此外,PAGE-RAG纳入严格知识边界控制,确保生成的回答基于可用证据,避免超出可访问知识范围的无支撑信息。实验表明,PAGE-RAG在提高检索效率和知识可靠性的同时,实现了有竞争力的答案质量,突出了投影感知图建模、自适应检索和明确知识边界控制对可信GraphRAG系统的重要性。源代码可在指定网址公开获取。
英文摘要
GraphRAG improves long-document question answering by introducing structured representations beyond conventional retrieval. However, automatically constructed graphs are inherently incomplete projections of source documents, and treating them as independent knowledge sources may lead to unreliable retrieval and generation. We propose PAGE-RAG, a projection-aware adaptive graph retrieval framework for reliable long-document question answering. PAGE-RAG views graph structures as semantic skeletons that organize and navigate document knowledge, rather than replacing the original knowledge source. Based on this perspective, PAGE-RAG introduces a task-adaptive retrieval routing strategy that dynamically selects appropriate retrieval behaviors according to query requirements. Furthermore, PAGE-RAG incorporates strict knowledge boundary control, ensuring that generated responses remain grounded within available evidence and abstaining from unsupported information beyond the accessible knowledge scope. Experiments demonstrate that PAGE-RAG achieves competitive answer quality while improving retrieval efficiency and knowledge reliability, highlighting the importance of projection-aware graph modeling, adaptive retrieval, and explicit knowledge boundary control for trustworthy GraphRAG systems. The source code is publicly available at https://github.com/CXY0112/PAGE-RAG.
Comments22 pages, 2 figures, and 3 tables. The source code is publicly available at https://github.com/CXY0112/PAGE-RAG