arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16814cs.AIcs.IR

我们能基于原子命题的图进行可解释的自然语言推理吗?

Can We Do Interpretable NLI with Graphs Based on Atomic Propositions?

  • Université Paris-Saclay(巴黎-萨克雷大学)
  • CNRS(法国国家科学研究中心)
  • Laboratoire Interdisciplinaire des Sciences du Numérique(数字科学跨学科实验室)

机构由 AI 辅助整理,请以论文原文为准。

Younes Boufouss, Luc Pommeret, Thomas Gerald, Patrick Paroubek, Sophie Rosset

AI总结:

本文提出一种完全基于图的NLI流程,将句子分解为原子命题并构建ConceptNet图,输入0.8B参数模型,在SNLI上达89.7%准确率,虽略低于文本模型,但揭示了可解释性的代价,且图与文本模态互补。

AI中文摘要:

虽然基于大型语言模型(LLM)的自然语言推理(NLI)系统实现了高准确率,但其决策过程缺乏可审计的结构。本文探讨是否可以使用仅基于可解释的、图结构的证据表示来进行NLI。我们引入了一个完全基于图的流程,其中分类器从不直接处理输入文本。相反,句子被分解为原子命题,通过约束解码转换为ConceptNet三元组,并表示为每对三个图:前提、假设和检索到的ConceptNet子图。这些图随后被输入到一个经过微调的0.8亿参数语言模型中。在SNLI数据集上,我们的流程达到了89.7%的准确率,仅比同等训练的基于文本的模型低1.9个百分点。在ANLI上,它在R2和R3轮次上匹配了RoBERTa-large的已发表性能(50%准确率),但在R1上落后16个百分点,导致与文本对应模型相比总体差距为9到14个百分点。我们将这一差距称为可解释性的代价,并证明其源于表示限制而非数据约束。消融研究进一步表明,图和文本是互补的:在SNLI上结合两种模态达到了92.1%的准确率。

英文摘要:

While Large Language Model (LLM)-based Natural Language Inference (NLI) systems achieve high accuracy, their decision-making processes lack auditable structures. This paper explores whether NLI can be performed using only interpretable, graph-based representations of evidence. We introduce a fully graph-based pipeline where the classifier never directly processes the input text. Instead, sentences are decomposed into atomic propositions, converted into ConceptNet triples via constrained decoding, and represented as three graphs per pair: premise, hypothesis, and a retrieved ConceptNet subgraph. These graphs are then fed into a fine-tuned 0.8-billion-parameter language model. On the SNLI dataset, our pipeline achieves 89.7% accuracy, just 1.9 points below an identically trained text-based model. On ANLI, it matches the published performance of RoBERTa-large on rounds R2 and R3 (48.0% vs. 48.9% and 44.9% vs. 44.4%) but trails by 16 points on R1, resulting in an overall gap of 9 to 14 points compared to its text counterpart. We term this gap the price of interpretability and demonstrate that it stems from representational limitations rather than data constraints. Ablation studies further reveal that graphs and text are complementary: combining both modalities achieves 92.1% accuracy on SNLI.

补充信息

↑