发表机构
Zhejiang University; National University of Singapore(浙江大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出DatalogBench基准,含136个文本到Datalog合成任务,评估六个LLM和两个编码智能体,发现精确匹配最高68.4%,智能体达83.8%,揭示递归推理与分解是当前挑战。
AI 中文摘要
Datalog支撑着诸如程序分析等推理任务,但其程序难以编写。现有的合成器可以自动化这一任务,但要求用户以输入输出示例的形式陈述其意图。大型语言模型(LLMs)提供了一种更自然的途径,即从自然语言问题直接进行文本到Datalog的合成,然而其效果尚未得到系统评估。我们提出了DatalogBench,一个包含136个文本到Datalog合成任务的基准,这些任务源自现有的基于Datalog的工件。合成程序通过在保留的输入上执行,与经过变异分析验证的预言机进行比较来评分。在六个LLM和四种提示配置下,精确匹配最高达到68.4%,而关系描述或输入输出示例仅产生适度且依赖模型的影响。在直接提示下,大多数失败发生在编译阶段,典型原因是模型发明了从未声明或类型不一致的辅助谓词。两个编码智能体最高达到83.8%,并几乎消除了所有此类失败,仅留下主要集中在递归任务中的语义错误。因此,DatalogBench将递归推理和分解识别为当前LLM和智能体面临的开放挑战,并为两者提供了可靠的、基于执行的度量标准。
英文摘要
Datalog underpins reasoning tasks such as program analysis, but its programs are hard to write. Existing synthesizers automate this task but require users to state their intent as input-output examples. Large language models (LLMs) suggest a more natural route, text-to-Datalog synthesis from a natural-language question, yet how well they do so has not been systematically evaluated. We present DatalogBench, a benchmark of 136 text-to-Datalog synthesis tasks curated from existing Datalog-based artifacts. Synthesized programs are graded by execution on held-out inputs against an oracle validated by mutation analysis. Across six LLMs and four prompting configurations, exact match peaks at 68.4%, and relation descriptions or an input-output example have only modest, model-dependent effects. Under direct prompting, most failures occur at compile time, typically because a model invents auxiliary predicates that it never declares or types consistently. Two coding agents reach up to 83.8% and eliminate nearly all such failures, leaving mostly semantic errors concentrated in recursive tasks. DatalogBench thus identifies recursive reasoning and decomposition as open challenges for current LLMs and agents, and offers a reliable, execution-grounded measure of both.
Comments33 pages