发表机构
Liber AI Research(Liber AI研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出研究者智能体式文本转SPARQL系统,在DBpedia上迭代优化后,其在2025年DBpedia验证集上总体准确率达0.22,发现瓶颈在谓词选择,建议基准采用多指标评分。
AI 中文摘要
将自然语言问题转换为可在大型知识图谱上执行的SPARQL查询,需要解决词汇歧义、将表层术语锚定到目标本体,同时生成在语法上有效且语义忠实的图模式。我们提出了一种智能体式的文本转SPARQL系统,它超越了静态工具使用智能体:一种研究者智能体,在每轮验证集推理后,会提出并测试对自身提示、规则和工具编排代码的修改。我们在DBpedia上实例化该循环,由低成本推理模型驱动生成了该智能体的9个连续版本,并将性能最佳的配置与两个更强的骨干模型一起部署。该研究得出三个观察结果:(i)自我改进快速收敛,随后在2025年DBpedia验证集上达到0.22的总体准确率;(ii)瓶颈始终在于基本图模式的谓词选择,而非SPARQL语法或修饰符;(iii)若干基准项似乎因DBpedia中的属性歧义而惩罚正确查询,这表明未来的文本转SPARQL基准应使用机器翻译和信息检索指标的组合进行评分。
英文摘要
Translating a natural-language question into a SPARQL query that can be executed against a large knowledge graph requires resolving lexical ambiguity, grounding surface terms in the target ontology, and producing graph patterns that are both syntactically valid and semantically faithful. We present an agentic text-to-SPARQL system that goes one step beyond static tool-using agents: a researcher agent that, after each round of inference on a validation set, proposes and tests changes to its own prompts, rules, and tool-orchestration code. We instantiate the loop on DBpedia, evolve nine successive versions of the agent driven by a low-cost reasoning model, and deploy the best-performing configuration with two stronger backbone models. The study yields three observations: (i) self-improvement converges quickly and then achieves 0.22 overall accuracy on the 2025 DBpedia validation set; (ii) the bottleneck is consistently in basic-graph-pattern predicate selection, not in SPARQL syntax or modifiers; and (iii) several benchmark items appear to penalise correct queries due to property ambiguity in DBpedia, suggesting that future Text-to-SPARQL benchmarks should be scored using a combination of machine translation and information retrieval metrics.
CommentsSecond International TEXT2SPARQL Challenge, co-located with Text2KG at ESWC 2026