arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GROUND:通过受监管的语义定义减少基于大语言模型的企业分析中的幻觉现象

GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions

Aravind Sasidharan Pillai

arXiv 2608.26157首次发表:更新:

发表机构

Cox Automotive Inc(考克斯汽车公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出GROUND框架,通过受监管语义层约束大语言模型生成的企业分析内容,在多模型多数据集测试中消除了所有评估类别的幻觉,保障了数据安全。

AI 中文摘要

针对企业数据仓库的自然语言分析日益重要,但生产环境中的应用受到幻觉指标、无效连接、错误粒度、不安全数据访问以及不支持的解释的限制。现有的文本转SQL系统通常通过数据库模式或检索到的文档来约束生成内容,而企业报告还需要受监管的业务语义:经批准的指标、维度、连接路径、筛选条件以及行级安全。本文介绍了GROUND(Governed Retrieval Over Unified Normalized Definitions,基于统一标准化定义的受监管检索),这是一个将大语言模型生成的分析约束在受监管语义层中的框架。GROUND提供经批准的定义,将用户意图绑定到受监管的指标和维度,并在执行前根据模式、指标、连接、粒度、筛选条件、安全和成本规则验证生成的SQL。若违反规则,它会重试或弃权(不执行)。在包含100个问题的合成企业报告基准测试中,GROUND与仅基于模式的文本转SQL、模式检索增强生成(RAG)以及仅基于语义的约束在同一模型下进行了比较。GROUND是唯一在所有六个评估类别中均未出现可测量幻觉的系统,而不受监管的系统在许多问题上违反了行级安全。具有精确指标定义但无访问策略的仅语义约束条件仍会泄露数据,表明监管不能仅靠指标保真度来替代。该发现通过独立手工编写的黄金标准在真实的美国国家公路交通安全管理局(NHTSA)车辆安全数据上得到复制,并在三个提供商的四个模型组成的对抗性测试集上进行了测试。GROUND的强制保障措施,尤其是筛选条件和行级安全,在每个模型上均保持零违反,而依赖判断的行为(例如拒绝未定义的指标)仍然容易出错。

英文摘要

Natural-language analytics over enterprise data warehouses is increasingly important, but production use is limited by hallucinated metrics, invalid joins, wrong grain, unsafe data access, and unsupported explanations. Existing text-to-SQL systems often ground generation in database schemas or retrieved documentation, while enterprise reporting also requires governed business semantics: approved metrics, dimensions, join paths, filters, and row-level security. This paper introduces GROUND, Governed Retrieval Over Unified Normalized Definitions, a framework that constrains LLM-generated analytics to a governed semantic layer. GROUND supplies approved definitions, binds user intent to governed metrics and dimensions, and validates generated SQL against schema, metric, join, grain, filter, security, and cost rules before execution. On violations, it retries or abstains. In a 100-question synthetic enterprise-reporting benchmark, GROUND is compared with direct schema-only text-to-SQL, schema-RAG, and semantic-only grounding under one shared model. GROUND is the only system free of measured hallucinations across all six evaluated categories, while ungoverned systems violate row-level security on many questions. A semantic-only condition with exact metric definitions but no access policy still leaks data, showing that governance cannot be replaced by metric fidelity alone. The findings are replicated on real U.S. NHTSA vehicle-safety data with independent hand-authored gold and tested on an adversarial set across four models from three providers. GROUND's enforced guarantees, especially filters and row-level security, hold with zero violations on every model, while judgment-dependent behaviors such as refusing undefined metrics remain fallible.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑