arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

针对哪个预言机认证?执行标签决定文本到SQL的保形弃权(不执行)的报告风险

Certified Against Which Oracle? Execution Labels Set the Reported Risk of Conformal Abstention for Text-to-SQL

Jiamiao Liu, Dewen Qiao, Yu Zhang, Xuetao Chen

arXiv 2609.25938首次发表:更新:

发表机构

Xinqiao Hospital; Army Medical University (Third Military Medical University)(新桥医院; 陆军军医大学(第三军医大学))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对文本到SQL的保形弃权(不执行)证书,其风险取决于校准所用预言机;用更严格的测试套件替换基准数据库会使报告风险显著上升,且专家标签下风险更高,因此证书应报告两个预言机,一致性得分应在未构建它的预言机下评估。

AI 中文摘要

针对文本到SQL的保形弃权(不执行)证书,其真实性取决于其校准所依据的正确性标签。从执行一致性中读取置信度的不确定性流程,从基准测试附带的一个数据库中获取这些标签,这个数据库是一个已知宽松的预言机。我们在Spider-Realistic上进行了预注册干预,将该数据库替换为基准的蒸馏多实例测试套件。在四个SQL专家检查点和两种分割方案中,这种替换使证书的留出风险比其自身标签报告的风险高出2.73到10.23个百分点。两个预言机都没有报告专家分配的风险。在两位SQL专家的盲标签下,以名义0.10校准的证书在两个检查点上分别携带20.0和17.2个百分点的风险。更严格的预言机在两个方向上都出错:它拒绝的大多数答案并未被判定为错误,而它接受的一些答案却被判定为错误。对其拒绝内容的AI分配普查发现,根据总体不同,四分之一到三分之一存在语义错误。它将其余大部分归因于问题说明不充分、合成实例或疑似参考查询缺陷,这一标记得到了预注册盲专家审计的支持。预言机还决定了置信度得分如何被评判。在16个组合中的16个中,每个执行一致性得分在构建其聚类的预言机的标签下看起来更好。在专家标签下,基于套件聚类而非附带数据库聚类构建这样的得分,在一个检查点上将其ROC曲线下面积(AUROC)提高了6.96个百分点,在另一个检查点上提高了1.53个百分点。在第二个检查点上,专家区间排除了套件标签报告的8.3个百分点。证书应同时报告两个预言机的结果,并且只有在基准经过审计后,预言机相对差异才能被解读为语义风险。一致性得分应在未构建它的预言机下进行评估。

英文摘要

A conformal abstention certificate for text-to-SQL is only as truthful as the correctness labels it is calibrated on. The uncertainty pipelines that read confidence off execution consistency take those labels from the single database a benchmark ships, an oracle known to be lenient. We run a preregistered intervention on Spider-Realistic, swapping that database for the benchmark's distilled multi-instance test suite. Across four SQL-specialist checkpoints and two split schemes, the swap raises the certificate's held-out risk 2.73 to 10.23 points above the risk its own labels report. Neither oracle reports the risk experts assign. Under blinded labels from two SQL experts, a certificate calibrated at a nominal 0.10 carries 20.0 and 17.2 points of risk on two checkpoints. The stricter oracle errs in both directions: most of the answers it rejects are not judged wrong, and some of those it accepts are. An AI-assigned census of what it rejects finds a semantic error in a quarter to a third of them, depending on the population. It attributes most of the rest to underspecified questions, synthetic instances or suspected reference-query defects, a flag supported by a preregistered blinded expert audit. The oracle also decides how a confidence score is judged. Every execution-consistency score looks better under the labels of the oracle that built its clusters, in 16 of 16 combinations. Under expert labels, building such a score on suite clusters instead of shipped-database clusters raises its area under the ROC curve (AUROC) by 6.96 points on one checkpoint and 1.53 on the other. On the second, the expert interval excludes the 8.3 points the suite labels report. A certificate should be reported with both oracles, and an oracle-relative difference read as semantic risk only after the benchmark is audited. A consistency score should be evaluated under an oracle that did not build it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑