SALUS:通过多智能体输出的弱监督实现自然语言到SQL基准的自动化审计
SALUS: Automated Auditing of NL-to-SQL Benchmarks through Weak Supervision of Multi-Agent Output
浏览论文内容
中文总结 AI 辅助
SALUS利用多智能体弱监督自动审计NL-to-SQL基准中的标注错误,在BIRD-Clean-xs上达到F1=0.9194,并估计BIRD和Spider的错误率分别约为37%和27%。
中文摘要 AI 辅助
自然语言到SQL(NL-to-SQL)基准是数据分析研究进展的基础,然而近期研究表明,广泛使用的基准包含大量标注错误。这些错误会悄无声息地破坏评估指标,惩罚正确的模型输出,并扭曲领域对最先进性能的理解。我们提出SALUS,一个自动检测NL-to-SQL基准中标注错误的系统。SALUS将基准审计构建为弱监督错误检测:由多个LLM智能体生成的SQL驱动一套互补的弱标注函数。通过将这个带噪声的投票矩阵输入生成式标签模型,我们在无需人工真值的情况下提取高置信度的训练样本。这些样本训练一个决策平面,将金标准SQL查询特征映射到每个智能体的可信度,使SALUS能够智能地将可靠性估计与原始判定融合,以实现严格的基准错误检测。我们在BIRD-Clean-xs上评估,该基准包含298个BIRD开发任务,并带有手动验证的正确性标签。SALUS达到F1=0.9194,显著优于最先进的基线方法。将SALUS应用于完整开发集,我们估计BIRD上的标注错误率约为37%,Spider上约为27%。
英文摘要
Natural language to SQL (NL-to-SQL) benchmarks are foundational to progress in data analysis research, yet recent work has shown that widely-used benchmarks contain significant annotation errors. These errors silently corrupt evaluation metrics, penalize correct model output, and distort the field's understanding of state-of-the-art performance. We present SALUS, a system that automatically detects annotation errors in NL-to-SQL benchmarks. SALUS frames benchmark auditing as a weakly supervised error detection: SQL generated by multiple LLM agents drive a suite of complementary weak-labeling functions. By passing this noisy vote matrix through a generative label model, we extract high-confidence training samples without requiring human ground truth. These samples train a decision plane that maps gold SQL query features to per-agent trustworthiness, allowing SALUS to intelligently fuse reliability estimates with raw verdicts for rigorous benchmark error detection. We evaluate on BIRD-Clean-xs, a benchmark of 298 BIRD development tasks with manually verified correctness labels. SALUS achieves F1 = 0.9194, significantly outperforming the state-of-the-art baselines. Applying SALUS to the full development sets, we estimate annotation error rates of approximately 37% on BIRD and 27% on Spider.
发表机构
- University of California, Irvine(加州大学欧文分校)
- Cornell University(康奈尔大学)
机构由 AI 辅助整理,请以论文原文为准。