AI 中文总结
该研究针对LLM分析智能体的评估缺陷,提出WarehouseReliabilityBench基准,开发规则门控的7B智能体QueryProof,其业务真值率优于直接提示的32B基线,且成本显著降低,假成功比例大幅下降。
AI 中文摘要
现有大语言模型(LLM)分析智能体的评估基于SQL语法准确率,但生产环境中的故障表现不同:存在具有两种有效业务定义的问题、数据仓库无法回答的问题、模式变更后废弃的列,以及执行成功却返回错误业务数值的查询。没有任何执行匹配指标能对这些情况打分。本文提出WarehouseReliabilityBench,这是针对两个合成数据仓库的400个固定任务,其中约一半的正确响应为澄清、弃权(不执行)或拒绝,且设置了固定分母,并采用预先注册的配对自举法,在数值确定前固定每个声明动词。QueryProof是一款7B智能体,它使用从语义层和物理目录衍生的规则来确定自身行为,并对每个答案施加确定性的执行后检查。在评估的80个合成测试任务中,QueryProof的业务真值率比直接提示的32B基线高+0.237(95%置信区间为[+0.112, +0.375]),且每个正确答案的成本低71.0%;与成本匹配的少样本基线相比,准确率提升依然存在,但成本差异不再显著。该研究是在对比系统而非模型规模:32B基线未获得任何辅助支持。返回答案中的假成功比例从0.754降至0.351,且在可回答任务中未返回任何错误数值(24个任务中0个),尽管有13个答案对应需要澄清或弃权(不执行)的问题。移除路由层后结果变化不大(0.562对比0.537),因此结果不依赖于升级机制。在验证集上调优的路由层在测试集上过度弃权(不执行),且拟合的置信模型性能逊于其替代的启发式方法。对模板家族而非任务进行重采样会使两个准确率区间扩大至包含零,因此效果的方向比其幅度更具支撑性。该提升与确定性层相关,但未进行组件消融实验。
英文摘要
LLM analytics agents are evaluated on SQL syntax accuracy, but production failures look different: questions with two valid business definitions, questions the warehouse cannot answer, deprecated columns after a schema change, and queries that execute successfully while returning the wrong business number. No execution-match metric can score them. This paper introduces WarehouseReliabilityBench, 400 frozen tasks over two synthetic warehouses in which roughly half the correct responses are a clarification, an abstention or a refusal, with pinned denominators and a pre-registered paired bootstrap fixing each claim verb before the numbers existed. QueryProof, a 7B agent, uses rules derived from a semantic layer and physical catalog to determine its behaviour, and gates every answer on deterministic post-execution checks. On an 80-task synthetic test split evaluated once, QueryProof outperforms a direct-prompted 32B baseline by +0.237 [+0.112, +0.375] Business Truth Rate at 71.0% lower cost per correct answer; against a cost-matched few-shot baseline the accuracy gain holds but the cost difference does not resolve. This compares systems rather than model sizes: the 32B baseline receives none of the scaffolding. False success falls from 0.754 to 0.351 of returned answers, and no wrong number was returned on an answerable task (0 of 24), though 13 answers went to questions requiring clarification or abstention. Removing the routing layer changes little (0.562 against 0.537), so the result does not depend on escalation. Routing tuned on validation over-abstains on test, and the fitted confidence model loses to the heuristic it replaced. Resampling template families rather than tasks widens both accuracy intervals to include zero, so the effect's direction is better supported than its magnitude. The gain tracks the deterministic layer, though no component ablation was run.
Comments20 pages, 7 figures, 11 tables. Benchmark, code, run records, pre-registration and one-command reproduction: https://github.com/k-w-lee/query_proof