arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AptMQL-Bench:从文本到SQL到文本到MQL——基于访问模式模式设计与数据保留迁移

AptMQL-Bench: From Text-to-SQL to Text-to-MQL via Access-Pattern Schema Design and Data-Preserving Migration

Hy Nguyen, Nabi Rezvani, Robin Vujanic

arXiv 2610.02770首次发表:更新:

发表机构

The University of Sydney; MongoDB Research(悉尼大学; MongoDB 研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有文本到SQL到文本到MQL转换的缺陷,提出基于访问模式模式设计与数据保留迁移的流水线,构建AptMQL-Bench基准,最强模型准确率仅57.38%,表明该任务仍具挑战性。

AI 中文摘要

MongoDB等文档数据库是现代应用的核心基础设施,而面向它们的自然语言接口——文本到MQL——能够让非专家用户无需掌握查询语言即可查询复杂的半结构化数据。该任务的进展依赖于高质量的基准测试,而最实用的获取方式是将现有的文本到SQL基准测试转换到文档设置中。不幸的是,现有的工作依赖于启发式方法进行机械转换:文档模式镜像了关系型外键图,每个查询也镜像其源SQL。因此,在我们的实验中,这些方法直接无法迁移21个BIRD数据库中的6个,在其他数据库上静默丢弃多达25.9%的行,并且产生的模式中,真实查询的运行速度随着数据规模扩大而慢了一个数量级以上。我们转而提出一种由编码代理驱动并辅以人工验证的转换流水线,该流水线根据预期的访问模式设计每个文档模式,并将查询重写为MongoDB原生形式。将其应用于BIRD,我们构建了一个基于访问模式的文本到MQL基准测试(AptMQL-Bench)。它包含21个面向文档的数据库、3186个自然语言请求及其关联的MQL查询——这些数据库从SQLite迁移而来,无数据丢失且扩展高效。最强的模型Claude Opus 4.5在无外部知识证据的情况下仅达到57.38%的准确率,在有外部知识证据时达到70.34%。这表明现实的文本到MQL生成仍然具有挑战性。

英文摘要

Document databases such as MongoDB are core infrastructure for modern applications, and natural-language interfaces to them---text-to-MQL---would let non-experts query complex, semi-structured data without mastering the query language. Progress on this task depends on high-quality benchmarks, which are most practically obtained by converting an existing text-to-SQL benchmark to the document setting. Unfortunately, existing efforts rely on heuristics for mechanical conversion: the document schema mirrors the relational foreign-key graph, and each query mirrors its source SQL. As a result in our experiments, these approaches fail to migrate 6 of 21 BIRD databases outright, silently drop up to 25.9\% of rows on others, and yield schemas whose ground-truth queries run over an order of magnitude slower as the data scales. We instead propose a conversion pipeline, driven by coding agents with human-in-the-loop verification, that designs each document schema from expected access patterns and rewrites queries to be MongoDB-native. Applying it to BIRD, we build an access-pattern-based text-to-MQL benchmark (AptMQL-Bench). It includes 21 document-oriented databases, 3,186 natural-language requests, and their associated MQL queries---whose databases are migrated from SQLite without data loss and scale efficiently. The strongest model, Claude Opus 4.5, achieves only 57.38\% accuracy without external knowledge evidence and 70.34\% with it. This indicates that realistic text-to-MQL generation remains challenging.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑