发表机构
Waseda University(早稻田大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究文本到SQL中列级访问控制问题,提出PCC-SQL系统,通过整合策略并应用逐令牌对数掩码,在单次解码中消除违规,在三个基准和开源模型上实现0%泄漏率与88.7%覆盖率,还评估了语义对齐。
AI 中文摘要
文本到SQL越来越多地部署在数据提供者和用户之间的信任边界上。这种部署必须平衡三个相互竞争的要求:策略合规性、答案覆盖范围和有限成本。现有方法通常根据查询提及的列来决定拒绝并随机执行。然而,查询是否合规不仅取决于出现哪些列,还取决于它们的使用方式,随机执行无法确定性地排除违规。我们将此要求形式化为语义使用上的列使用策略:输出、过滤条件和聚合参数。我们通过将每个角色与解码器跟踪的语法产生式对齐来整合策略。所得系统PCC-SQL应用逐令牌对数掩码,在单次解码过程中确定性地消除支持的SQL片段上的单查询列使用违规。在三个基准和三个开源模型上,PCC-SQL在Spider-CU上实现了0%的泄漏率和高达88.7%的覆盖率,同时保持在直接提示的+10%令牌范围内。我们还评估了与执行准确性的语义对齐。
英文摘要
Text-to-SQL is increasingly deployed across trust boundaries between data providers and users. Such deployment must balance three competing requirements: policy compliance, answer coverage, and bounded cost. Existing approaches typically decide refusal based on which columns a query mentions and enforce it stochastically. Whether a query is compliant, however, depends not only on which columns appear but on how they are used, and stochastic enforcement cannot deterministically rule out violations. We formalize this requirement as a column-use policy over semantic use: output, filter condition, and aggregation argument. We integrate the policy by aligning each role with grammar productions tracked by the decoder. The resulting system, PCC-SQL, applies a per-token logits mask that deterministically eliminates single-query column-use violations on the supported SQL fragment in a single decoding pass. Across three benchmarks and three open-source models, PCC-SQL achieves 0% Leakage Rate and Coverage up to 88.7% on Spider-CU, while staying within +10% tokens of direct prompting. We additionally assess semantic alignment with execution accuracy.
CommentsAACL-IJCNLP 2026 Main