arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.14494cs.AIcs.IRcs.LG

SAGA:用于智能文本到SPARQL生成的模式感知基础

SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation

Yiming Zhang, Koji Tsuda

首次发表
浏览论文内容

中文总结 AI 辅助

研究复杂知识库问答中语义解析范式,针对现有智能体类型盲基础问题,提出无训练框架SAGA,通过模式约束基础操作,维护类型状态,过滤候选属性,以紧凑格式呈现图模式,在多基准设置上取得优异成绩。

中文摘要 AI 辅助

复杂知识库问答(KBQA)通常通过特定问题子图上的信息检索或语义解析为可执行逻辑形式来解决。本文研究后一种范式。近期大语言模型智能体使语义解析具有交互性,但其知识库基础的质量依赖于交互工具。现有智能体主要通过词汇相关性和实例级观察检索或修剪候选属性,存在类型盲基础问题。本文提出SAGA,一个无训练框架,将属性探索转化为模式约束基础操作。SAGA维护持久双向类型状态,在构建时过滤已知不兼容属性候选,以紧凑模式注释格式呈现剩余图模式,并通过经验和跟踪局部证据宽松处理缺失模式信息。在九个基准设置上,SAGA在所有九个设置中实现了最高F1,在八个设置中实现了最高精确匹配准确率,同时减少了所有报告的维基数据设置中的空结果查询。

英文摘要

Complex knowledge base question answering (KBQA) is commonly approached through either information retrieval over a question-specific subgraph or semantic parsing into an executable logical form. We study the latter paradigm. Recent large language model agents make semantic parsing interactive: they alternate between reasoning, querying the knowledge base, and extending a partial SPARQL query. This interleaving reduces reliance on one-shot generation, but makes the quality of \emph{KB grounding} depend on what the interaction tools expose. Existing agents retrieve or prune candidate properties mainly through lexical relevance and instance-level observations, without systematically conditioning on entity types, property domains and ranges, or the expected answer type. We call this failure mode \emph{type-blind grounding}. It enlarges the grounding search space and often produces plausible-looking but semantically incompatible triple patterns that execute to empty results. We propose SAGA (\underline{S}chema-\underline{A}ware \underline{G}rounding for \underline{A}gentic Text-to-SPARQL Generation), a training-free framework that turns property exploration into a schema-constrained grounding operation. SAGA maintains a persistent bidirectional type state, filters known-incompatible property candidates at construction time, presents the remaining graph patterns in a compact schema-annotated format, and handles missing schema information permissively through empirical and trace-local evidence. Across nine benchmark settings over Wikidata and Freebase, SAGA achieves the highest F1 on all nine settings and the highest exact-match accuracy on eight, while reducing empty-result queries across all reported Wikidata settings.

↑