发表机构
Hunan University(湖南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出TQTS-Bench多语法基准,含6125个QA对覆盖97个TSDBs,评估LLM文本到查询能力,发现最佳模型准确率仅48.98%,远低于人类的87.34%,揭示语法异构等挑战。
AI 中文摘要
大型语言模型(LLMs)显著推进了关系数据库上的自然语言查询,但其查询时间序列数据库(TSDBs)的能力在很大程度上仍未得到评估。现有基准未能充分捕捉TSDBs固有的非统一查询语法、多样化的应用领域以及独特的时间特定查询意图。为弥补这一空白,我们引入了TQTS-BENCH,一个用于评估TSDBs上文本到查询能力的多语法基准。TQTS-BENCH包含6,125个高质量问答(QA)对,涵盖97个TSDBs、23种不同的查询语法、22个应用领域以及4种时间特定查询意图。它通过以人为中心的AI辅助工作流构建,所有QA对均由领域专家仔细审查和修订,以确保质量和正确性。对先进LLMs和最先进的文本到查询方法的广泛评估揭示了查询TSDBs的挑战。即使评估中表现最好的模型Claude-Opus-5,也仅达到48.98%的执行准确率,而人类达到87.34%。错误分析表明,这一性能差距主要源于不同TSDBs之间的异构查询语法、对时间特定意图的误解以及错误的模式链接。这些发现凸显了缩小当前LLM能力与现实应用中TSDB查询需求之间差距的新机遇。该基准可在以下网址获取:this https URL。
英文摘要
Large language models (LLMs) have significantly advanced natural language querying over relational databases, yet their ability to query time-series databases (TSDBs) remains largely unassessed. Existing benchmarks fail to adequately capture the non-unified query syntaxes, diverse application domains, and unique time-specific query intents inherent to TSDBs. To address this gap, we introduce TQTS-BENCH, a multi-syntax benchmark for evaluating text-to-query capabilities over TSDBs. TQTS-BENCH contains 6,125 high-quality question-answering (QA) pairs spanning 97 TSDBs, 23 distinct query syntaxes, 22 application domains, and 4 types of time-specific query intents. It is constructed through a human-centric AI-assisted workflow, where all QA pairs are carefully reviewed and revised by domain experts to ensure quality and correctness. Extensive evaluations of advanced LLMs and state-of-the-art text-to-query methods reveal challenges in querying TSDBs. Even the best-performing model evaluated, Claude-Opus-5, achieves only 48.98% execution accuracy, while humans reach 87.34%. Error analysis reveals that this performance gap mainly stems from the heterogeneous query syntaxes across different TSDBs, misinterpretation of time-specific intents, and incorrect schema linking. These findings highlight new opportunities to narrow the gap between current LLM capabilities and the requirements of TSDB queries in real-world applications. The benchmark is available at: https://anonymous.4open.science/r/TQTS-Bench-00CD.