ScenarioBench:面向 Text-to-SQL 与 RAG 的基于追踪的合规性评估
ScenarioBench: Trace-Grounded Compliance Evaluation for Text-to-SQL and RAG
- Faculty of Business and Information Technology(商业与信息技术学院)
- Ontario Tech University(安大略技术大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出 ScenarioBench 合规评估基准,通过禁止预看的金标准追踪与条款级证据,端到端评测 Text-to-SQL 和 RAG 的决策、检索、SQL 及解释质量。
AI中文摘要:
ScenarioBench 是一个以政策为依据、具备追踪感知能力的基准,用于在合规情境中评估 Text-to-SQL 与检索增强生成。每个 YAML 场景都包含一个禁止预看的金标准包,其中有预期决策、最小见证追踪、适用条款集和规范 SQL,从而能够对系统的决策内容及其原因进行端到端评分。系统必须使用同一政策正典中的条款 ID 来证明输出的合理性,使解释可证伪且可供审计。评估器报告决策准确率、追踪质量(完整性、正确性、顺序)、检索效果、通过结果集等价性判定的 SQL 正确性、政策覆盖率、延迟以及解释幻觉率。归一化的 Scenario Difficulty Index(SDI)及其预算化变体 SDI-R 在汇总结果时考虑检索难度和时间。与以往的 Text-to-SQL 或 KILT/RAG 基准相比,ScenarioBench 在严格依据和禁止预看规则下将每项决策与条款级证据绑定,使收益转向明确时间预算下的论证质量。
英文摘要:
ScenarioBench is a policy-grounded, trace-aware benchmark for evaluating Text-to-SQL and retrieval-augmented generation in compliance contexts. Each YAML scenario includes a no-peek gold-standard package with the expected decision, a minimal witness trace, the governing clause set, and the canonical SQL, enabling end-to-end scoring of both what a system decides and why. Systems must justify outputs using clause IDs from the same policy canon, making explanations falsifiable and audit-ready. The evaluator reports decision accuracy, trace quality (completeness, correctness, order), retrieval effectiveness, SQL correctness via result-set equivalence, policy coverage, latency, and an explanation-hallucination rate. A normalized Scenario Difficulty Index (SDI) and a budgeted variant (SDI-R) aggregate results while accounting for retrieval difficulty and time. Compared with prior Text-to-SQL or KILT/RAG benchmarks, ScenarioBench ties each decision to clause-level evidence under strict grounding and no-peek rules, shifting gains toward justification quality under explicit time budgets.