arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自我演示在增强LLM模式-本体映射中的惊人有效性

Surprising Effectiveness of Self-Demonstrations in Enhancing Schema-Ontology Mapping with LLMs

Siddhesh Thombre, Manasi Patwardhan, Sunita Sarawagi

arXiv 2609.13776首次发表:更新:

发表机构

TCS Research; IIT Bombay(塔塔咨询服务研究院; 印度理工学院孟买分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种结合神经符号任务分解与模式引导自我演示的方法,用于LLM模式-本体映射,在RODI基准上实现最先进性能,F1提升25个百分点。

AI 中文摘要

将异构关系数据库集成到集中式本体中,仍然是企业知识表示中的一个持久挑战,主要源于语义异构性、隐晦的模式命名、缺失的元数据,以及关系模式与本体模型之间的抽象鸿沟。尽管大型语言模型(LLMs)提供了强大的语义推理能力,但我们表明,通过一次性提示或朴素的多阶段流水线直接应用它们,会导致模式-本体映射性能不佳。本文提出了一种自我演示驱动的方法,结合了神经符号任务分解与一种自动生成模式引导、依赖感知演示的新机制,以应对这一集成挑战。我们的方法包含两个关键策略,以实现相对于现有基于LLM的模式集成方法的显著准确性提升:(i)将任务神经符号地分解为级联子任务,其中符号约束构建搜索空间,LLMs在每个聚焦子任务内执行语义推理;(ii)由领域无关模式引导的自我生成演示,用于监督每个子任务。在RODI基准中三个最具挑战性场景上的实验表明,我们的方法达到了最先进的性能,大幅超越(F1提升25个百分点)了传统的模式到本体映射技术以及近期基于LLM的模式到本体和模式匹配方法。消融研究进一步揭示了模式引导的自我演示的显著益处,以及神经符号任务分解的互补优势。

英文摘要

Integrating heterogeneous relational databases into a centralized ontology remains a persistent challenge in enterprise knowledge representation, primarily due to semantic heterogeneity, cryptic schema naming, missing metadata, and the abstraction gap between relational schemas and ontological models. Although large language models (LLMs) offer strong semantic reasoning capabilities, we show that directly applying them through one-shot prompting or naive multi-stage pipelines leads to poor performance for schema-ontology mapping. This paper presents a self-demonstration-driven approach that combines a neuro-symbolic task decomposition with a novel mechanism for automatically generating pattern-guided, dependency-aware demonstrations to address this integration challenge. Our approach incorporates two key strategies to achieve substantial accuracy gains over existing LLM-based schema integration methods: (i) a neuro-symbolic decomposition of the task into cascaded sub-tasks, where symbolic constraints structure the search space and LLMs perform semantic reasoning within each focused sub-task, and (ii) self-generated demonstrations guided by domain-agnostic patterns to supervise each sub-task. Experiments on three of the most challenging scenarios from the RODI benchmark show that our approach achieves state-of-the-art performance, substantially outperforming (25 percentage points F1 improvements) both traditional schema-to-ontology mapping techniques and recent LLM-based schema-to-ontology and schema matching approaches. Ablation studies further reveal the significant benefits of pattern-guided self-demonstrations and the complementary benefits of neuro-symbolic task decomposition.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑