TraceSQL:无参考文本到SQL验证的可追踪答案性估计
TraceSQL: Traceable Answerability Estimation for Reference-Free Text-to-SQL Verification
- Oracle Corporation(甲骨文公司)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对无参考文本到SQL验证的可追踪性不足问题,提出基于67种诊断特征的轻量模型TraceSQL,在BIRD数据集上优于基线模型,可提供可追溯的预测证据。
AI中文摘要:
文本到SQL系统通常使用真实SQL查询或参考执行结果进行评估,但在实际部署的推理阶段无法获得此类监督,这就产生了一个关键的验证问题:仅给定用户问题、数据库上下文和生成的SQL,系统能否估计生成的查询是否可能正确回答该问题?近期方法使用大语言模型(LLM)作为评判者或专用智能体来检查生成的SQL,但其决策难以追踪;结果奖励模型(ORM)通过从带执行标签的候选SQL中学习并为未见过的查询分配正确性分数来解决这一问题,但仍无法为每次验证提供足够的信号透明度。为解决这一局限,我们提出TraceSQL,这是一种基于明确诊断特征构建的轻量且可追踪的验证模型。TraceSQL结合了67种特征,涵盖问题歧义性、问题需求、问题-模式-SQL一致性、SQL结构以及意图对齐,这些信号可用于检查哪些因素影响每次预测,并将决策追溯到诊断证据。在BIRD开发数据库上,TraceSQL在相同的生成SQL评估中达到了66.47%的F1值和64.48%的ROC-AUC,而基线模型GradeSQL-7B ORM的F1值为61.87%、ROC-AUC为58.26%。特征归因进一步表明,该模型同时依赖语义基础和确定性SQL结构信号。这些结果表明,SQL验证可通过轻量的学习模型完成,同时保留特征级证据以检查和诊断其预测。
英文摘要:
Text-to-SQL systems are commonly evaluated using ground-truth SQL queries or reference execution results, but such supervision is unavailable at inference time in real-world deployments. This creates a critical verification problem: given only a user question, database context, and generated SQL, can a system estimate whether the generated query is likely to correctly answer the question? Recent approaches use LLMs as judge or specialized agents to inspect generated SQL, but their decisions can be difficult to trace. Outcome Reward Models (ORMs) address this by learning from execution-labeled candidate SQLs and assigning correctness scores to unseen queries, yet they still provide limited visibility into the signals behind each verification. To address this limitation, we propose TraceSQL, a lightweight and traceable verification model built on explicit diagnostic features. TraceSQL combines 67 features capturing question ambiguity, question requirements, question-schema-SQL consistency, SQL structure, and intent alignment. These signals remain available for examining which factors influence each prediction and for tracing decisions back to diagnostic evidence. On BIRD development databases, TraceSQL achieves 66.47% F1 and 64.48% ROC-AUC, compared with 61.87% F1 and 58.26% ROC-AUC for the GradeSQL-7B ORM baseline on the same generated-SQL evaluation. Feature attribution further shows that the model relies on both semantic grounding and deterministic SQL-structure signals. These results show that SQL verification can be performed with a lightweight learned model while retaining feature-level evidence for inspecting and diagnosing its predictions.