arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06331cs.DBcs.AI

Tytan:从关系数据交互式神经符号构建分析语义模式

Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data

Donna Hooshmand, Shubham Shahi, Cameron Barrie, Abhratanu Dutta, Marko Sterbentz, Harper Pack, Kristian J. Hammond

首次发表
浏览论文内容

中文总结 AI 辅助

TYTAN是从关系数据自动构建分析语义模式的交互式神经符号系统,在多数据库评估中实现了100%覆盖率、检索正确性及高语义角色一致性,盲测表现优异。

中文摘要 AI 辅助

从自然语言查询界面到自动报告生成,数据分析工具需要对数据的描述:数据包含的现实实体、哪些列作为度量或标识符,以及表如何连接为分析单元。如今,这种语义层通常是手动编写的,这是一个知识获取瓶颈,限制了分析系统的可扩展性,使非技术用户依赖专家,且本身容易出错。我们提出TYTAN,这是一个从关系数据库(以及用户提供的简短描述)自动构建分析语义模式的系统。TYTAN将数据库的符号分析与基于大语言模型(LLM)的语义推理相结合,用于实体提议、角色分配和命名。当证据使决策模糊时,TYTAN会向用户提出针对性的自然语言问题。我们在涵盖现实世界和基准领域的8个数据库上,从定义语义模式功能效用的三个维度评估TYTAN:(i)覆盖率,是否捕获了所有重要实体和特征;(ii)检索正确性,语义模式的指令是否实际访问到数据;(iii)表征准确性,语义类型是否正确。在7个参考领域中,TYTAN覆盖了专家修正后的参考语义模式的所有实体、属性和可聚合特征(100%覆盖率)。此外,其100%的检索指令执行正确(1678个自生成声明中,1678个正确),且语义角色与参考匹配属性的一致性为92%-100%。检查底层数据后发现,小部分分歧出在参考语义模式,而非TYTAN。在一个保留的盲测(一个包含10个表且无声明键的实时数据库)中,TYTAN恢复了完整的实体结构及已验证的键,满足了5位独立盲标注者100%可满足的预期。

英文摘要

From natural-language query interfaces to automated report generation, data analysis tools need a description of the data: the real-world entities it contains, which columns function as measures or identifiers, and how tables connect into units of analysis. Today, this semantic layer is usually written by hand. This is a knowledge-acquisition bottleneck that limits the scalability of analytic systems, keeps non-technical users dependent on experts, and is itself error-prone. We present TYTAN, a system for automatically constructing an analytic semantic schema from a relational database and, when available, a short user-provided description. TYTAN combines symbolic analysis of the database with LLM-based semantic inference for entity proposal, role assignment, and naming. When the evidence leaves a decision ambiguous, TYTAN asks the user a targeted natural-language question. We evaluate TYTAN on eight databases spanning real-world and benchmark domains along the three axes that define a schema's functional utility: (i) coverage, are all important entities and features captured?; (ii) retrieval correctness, do the schema's instructions actually reach the data; and (iii) characterization accuracy, are semantic types correct? Across the seven reference domains, TYTAN reaches every entity, attribute, and aggregable feature of the expert-corrected reference schemas (100% coverage). Additionally, 100% of its retrieval instructions execute correctly (1,678 of 1,678 self-generated claims), and semantic roles agree with the reference on 92-100% of matched attributes. Checking the underlying data showed the small disagreement is in the reference, not in TYTAN. On a held-out blind test (a live, ten-table database with no declared keys), TYTAN recovers the full entity structure with verified keys and satisfies 100% of the satisfiable expectations of five independent blind annotators.

发表机构

  • Northwestern University(西北大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑