Tytan:从关系数据交互式神经符号构建分析语义模式
Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data
浏览论文内容
中文总结 AI 辅助
TYTAN是从关系数据自动构建分析语义模式的交互式神经符号系统,在多数据库评估中实现了100%覆盖率、检索正确性及高语义角色一致性,盲测表现优异。
中文摘要 AI 辅助
从自然语言查询界面到自动报告生成,数据分析工具需要对数据的描述:数据包含的现实实体、哪些列作为度量或标识符,以及表如何连接为分析单元。如今,这种语义层通常是手动编写的,这是一个知识获取瓶颈,限制了分析系统的可扩展性,使非技术用户依赖专家,且本身容易出错。我们提出TYTAN,这是一个从关系数据库(以及用户提供的简短描述)自动构建分析语义模式的系统。TYTAN将数据库的符号分析与基于大语言模型(LLM)的语义推理相结合,用于实体提议、角色分配和命名。当证据使决策模糊时,TYTAN会向用户提出针对性的自然语言问题。我们在涵盖现实世界和基准领域的8个数据库上,从定义语义模式功能效用的三个维度评估TYTAN:(i)覆盖率,是否捕获了所有重要实体和特征;(ii)检索正确性,语义模式的指令是否实际访问到数据;(iii)表征准确性,语义类型是否正确。在7个参考领域中,TYTAN覆盖了专家修正后的参考语义模式的所有实体、属性和可聚合特征(100%覆盖率)。此外,其100%的检索指令执行正确(1678个自生成声明中,1678个正确),且语义角色与参考匹配属性的一致性为92%-100%。检查底层数据后发现,小部分分歧出在参考语义模式,而非TYTAN。在一个保留的盲测(一个包含10个表且无声明键的实时数据库)中,TYTAN恢复了完整的实体结构及已验证的键,满足了5位独立盲标注者100%可满足的预期。
英文摘要
From natural-language query interfaces to automated report generation, data analysis tools need a description of the data: the real-world entities it contains, which columns function as measures or identifiers, and how tables connect into units of analysis. Today, this semantic layer is usually written by hand. This is a knowledge-acquisition bottleneck that limits the scalability of analytic systems, keeps non-technical users dependent on experts, and is itself error-prone. We present TYTAN, a system for automatically constructing an analytic semantic schema from a relational database and, when available, a short user-provided description. TYTAN combines symbolic analysis of the database with LLM-based semantic inference for entity proposal, role assignment, and naming. When the evidence leaves a decision ambiguous, TYTAN asks the user a targeted natural-language question. We evaluate TYTAN on eight databases spanning real-world and benchmark domains along the three axes that define a schema's functional utility: (i) coverage, are all important entities and features captured?; (ii) retrieval correctness, do the schema's instructions actually reach the data; and (iii) characterization accuracy, are semantic types correct? Across the seven reference domains, TYTAN reaches every entity, attribute, and aggregable feature of the expert-corrected reference schemas (100% coverage). Additionally, 100% of its retrieval instructions execute correctly (1,678 of 1,678 self-generated claims), and semantic roles agree with the reference on 92-100% of matched attributes. Checking the underlying data showed the small disagreement is in the reference, not in TYTAN. On a held-out blind test (a live, ten-table database with no declared keys), TYTAN recovers the full entity structure with verified keys and satisfies 100% of the satisfiable expectations of five independent blind annotators.
发表机构
- Northwestern University(西北大学)
机构由 AI 辅助整理,请以论文原文为准。