发表机构
Centro de Investigación en Ciencias, Universidad Autónoma del Estado de Morelos; Centro Interdisciplinario de Investigación en Humanidades, Universidad Autónoma del Estado de Morelos(莫雷洛斯州自治大学科学研究中心; 莫雷洛斯州自治大学跨学科人文研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出可解释混合模型SLITE,结合结构-关系与分布-信息特征,在SICK上达83%准确率,优于IsoLex且接近RoBERTa,验证了混合方法的科学价值。
AI 中文摘要
尽管神经模型在自然语言处理中取得了成功,但其黑箱特性限制了可解释性,并掩盖了其预测背后的语言学现象。我们提出了SLITE,一种用于识别文本蕴含的可解释混合模型,该模型整合了两个互补的语义分析层:一个基于组合实体间语义兼容性与不兼容性的结构-关系层,以及一个基于前提和假设的嵌入表示之间信息变化的结构化模式的分布-信息层。我们提出了17个特征,这些特征结合了实体级语义关系、极性敏感的词汇匹配,以及基于相似度矩阵的语义子表示的比对度量,包括基于熵和转移熵的度量。在这些特征上训练的逻辑回归模型在SICK三分类上达到了83%的准确率,在SICK-CE上达到了96%的准确率,比IsoLex高出4个百分点,并且与RoBERTa的差距在2个百分点以内,而计算复杂度仅为后者的一小部分。消融研究和SHAP分析证实,结构-关系特征是分类的主要驱动力,而分布-信息特征提供了必要的补充贡献,特别是在检测中立和矛盾方面。我们的结果表明,进一步探索混合方法是巨型神经架构的一种可行且科学高效的替代方案,我们希望这些结果能加强语言学理论与推理计算建模之间的对话。
英文摘要
Despite the success of neural models in natural language processing, their black-box nature limits interpretability and conceals the linguistic phenomena underlying their predictions. We present SLITE, an explainable hybrid model for Recognizing Textual Entailment that integrates two complementary layers of semantic analysis: a structural-relational layer, based on semantic compatibility and incompatibility between compositional entities, and a distributional-informational layer, based on structured patterns of information change between embedding-based representations of the premise and the hypothesis. We propose 17 features that combine entity-level semantic relations, polarity-sensitive lexical matching, and alignment measures over semantic sub-representations of the similarity matrix, including measures based on entropy and transfer entropy. A logistic regression trained on these features achieves an accuracy of 83% on three-class SICK and 96% on SICK-CE, outperforming IsoLex by 4 percentage points and falling within 2 percentage points of RoBERTa with a fraction of its computational complexity. Ablation studies and SHAP analysis confirm that structural-relational features are the primary drivers of classification, while distributional-informational features provide essential complementary contributions, particularly for detecting neutrality and contradiction. Our results demonstrate that further exploration of hybrid approaches is a viable and scientifically productive alternative to massive neural architectures, and we hope they will strengthen the dialogue between linguistic theory and computational modeling of inference
Comments38 pages, 5 figures, 8 tables