arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于角色的访问控制下的文本到SQL基准测试

Benchmarking Text-to-SQL under Role-Based Access Control

Yang Fei, Yangfan Jiang, Yin Yang, Xiaokui Xiao

arXiv 2607.22115首次发表:更新:

发表机构

National University of Singapore; Hamad Bin Khalifa University(新加坡国立大学; 哈马德·本·卡西姆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于角色的访问控制下文本到SQL的基准测试问题,提出含现实RBAC约束的综合基准测试框架,通过大语言模型辅助工作流程及人工审核等增强基准,应用该框架研究发现部分高分系统在有访问约束时性能因RBAC违规而大幅下降。

AI 中文摘要

给定数据库S和自然语言问题Q,文本到SQL系统旨在生成一个SQL查询,该查询在针对S执行时能正确回答Q。当前流行的文本到SQL基准大多假设对S可无限制访问,然而实际中用户访问常受限,如通过基于角色的访问控制(RBAC)策略。这导致基准测试结果与实际性能可能脱节。为此,我们提出一个具有现实RBAC约束的综合文本到SQL基准测试框架,其具有由大语言模型辅助的工作流程,通过合理的用户角色和访问策略增强现有基准。我们将角色合成问题表述为对数据库模式的结构化推理过程,大语言模型先从模式推断应用上下文,再得出与该上下文一致的角色职责和访问范围,此过程由人工参与的质量控制审核。该框架还包含评估指标,能识别特定于RBAC的失败模式,并区分SQL效用与访问控制合规性。我们将框架应用于多个广泛使用的基准,并对先进的文本到SQL系统进行系统实证研究。结果表明,许多在无限制设置下具有高基准分数的解决方案(尤其是开放权重的大语言模型),一旦存在访问约束,由于频繁违反RBAC,性能会急剧下降。

英文摘要

Given a database S and a natural language question Q, text-to-SQL systems aim to generate an SQL query that correctly answers Q when executed against S. Currently, popular text-to-SQL benchmarks mostly assume unrestricted access to S; in practice, however, user access is often restricted, e.g., through role-based access control (RBAC) policies. This leads to a potential disconnect between benchmarking results and real-world performance: an LLM with high benchmark scores might perform poorly in an access-controlled environment, by frequently violating RBAC, or rejecting a query q that could be answered with only permitted data in S. Motivated by this, we present a comprehensive text-to-SQL benchmarking framework with realistic RBAC constraints, which features an LLM-assisted workflow that augments existing text-to-SQL benchmarks with plausible user roles and access policies. To do so, we formulate the problem of role synthesis as a structured reasoning process over the database schema, in which the LLM first infers the application context from the schema, and then derives role responsibilities and access scopes consistent with this context. This process is audited by human-in-the-loop quality control, in which domain experts perform metric-guided screening on the generated roles. Besides the augmented dataset, the proposed framework also contains evaluation metrics that identify RBAC-specific failure modes, and disentangle SQL utility from access-control compliance. We apply the proposed framework to several widely-used benchmarks, and conduct a systematic empirical study of state-of-the-art text-to-SQL systems. The results show that many solutions (especially open-weight LLMs) with high benchmarking scores under an unrestricted setting suffer sharp performance degradation once access constraints are in place, due to frequent RBAC violations.

CommentsAccepted at ACM SIGMOD 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑