发表机构
Kellogg School of Management; Northwestern University; MIT(凯洛格商学院; 西北大学; 麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出用于论证分析的新框架,用封装论证逻辑语义的逻辑嵌入替代传统上下文词嵌入,基于RKHS理论定义,在分类任务中性能优于多数标准嵌入方法。
AI 中文摘要
我们提出了一种面向机器学习的论证分析任务的新框架。我们的方案是将大多数NLP任务中使用的传统上下文词嵌入替换为逻辑嵌入,这是一种直接利用论证结构的替代编码。本质上,逻辑嵌入封装了论证的逻辑语义,能更好地表示其含义。支撑这些嵌入的是一种基于数理逻辑的相似度度量,它提供了透明的邻近概念,且保证满足当前基于余弦相似度的上下文词嵌入无法确保的若干理想理论性质。该相似度度量在论证集合上诱导出一个半正定核,使我们能利用再生核希尔伯特空间(RKHS)理论唯一地定义逻辑嵌入。此外,我们证明这种编码是最优的,即过程中不会丢失任何逻辑信息。与其他RKHS应用一样,逻辑嵌入可用于众多有监督和无监督任务。我们提供了该方法的实现,旨在将其与文献基准进行测试。此外,我们在一项分类任务中证明,逻辑嵌入的性能优于大多数标准嵌入方法。
英文摘要
We propose a new framework for machine-learning-oriented argument analysis tasks. Our proposal involves replacing traditional contextualized word embeddings used in most NLP tasks with logical embeddings, an alternative encoding that directly exploits argumentation structures. In essence, logical embeddings encapsulate the logical semantics of an argument, allowing for a better representation of its meaning. Supporting these embeddings is a mathematical logic-based similarity measure that offers a transparent notion of proximity and is guaranteed to satisfy several desirable theoretical properties that current cosine similarity-based contextualized word embeddings cannot assure. This similarity measure induces a positive semi-definite kernel on the set of arguments, enabling us to uniquely define logical embeddings using the theory of Reproducing Kernel Hilbert Spaces (RKHS). Moreover, we prove that this encoding is optimal, in the sense that no logical information is lost in the process. As with other RKHS applications, logical embeddings can be used in numerous supervised and unsupervised tasks. We provide an implementation of the method and aim to test it against literature benchmarks. Additionally, we demonstrate that logical embeddings outperform most standard embedding methods on a classification task.