arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18961cs.AI

从依存关系到组合性:通过组合范畴语法对大语言模型输出进行神经符号提升

From Dependency to Compositionality: A Neurosymbolic Lifting of LLM Outputs via Combinatory Categorial Grammar

Remo Pareschi

首次发表
浏览论文内容

中文总结 AI 辅助

研究基于自回归生成与CCG增量处理模型的对齐,提出神经符号框架提升LLM输出为类型化组合推导,可扩展到多种形式语言,支持两层检查,还概述了同步LLM - CCG耦合方向。

中文摘要 AI 辅助

大语言模型(LLMs)通过根据前缀逐步预测下一个 token 来生成流畅文本。生成传统的批评者认为此类系统缺乏真正的语法;依存语法视角的有影响力的回应认为,LLM 的行为可以通过逐词构建的局部头部依存结构很好地描述。我们认为一个更敏锐的观察被忽视了:自回归生成的前缀驱动、类型完成动态与组合范畴语法(CCG)最初设计支持的增量处理模型紧密对齐。在此基础上,我们提出了一个神经符号框架,其中 LLM 输出被提升为类型化的组合推导——并非声称 LLMs 在内部实现了 CCG,而是它们的输出允许进行有原则的、增量的和可审计的 CCG 重建。这带来两个结果。首先,通过柯里 - 霍华德对应,这种提升不仅适用于自然语言,还适用于 LLMs 生成的形式语言,如 Solidity 等编程语言、OWL 和 SQL 等描述逻辑和查询语言,类型系统可变而架构固定。其次,这种提升支持两层检查:直接捕捉结构故障的组合层,以及根据外部知识源检查提升结构的内容层,从而能够尽早标记幻觉内容。该解释因此要求生产者具有的不是认知而是前缀驱动的生成配置文件。我们最后概述了同步 LLM - CCG 耦合,作为该框架开启的一个方向。

英文摘要

Large language models (LLMs) generate fluent text by incrementally predicting the next token from a prefix. Critics in the generative tradition argue that such systems lack genuine grammar; influential replies from the dependency-grammar perspective hold that LLM behavior is well described by local head-dependent structure built word by word. We argue that a sharper observation has been overlooked: the prefix-driven, type-completing dynamics of autoregressive generation align closely with the incremental processing model that Combinatory Categorial Grammar (CCG) was originally designed to support. On this basis we propose a neurosymbolic framework in which LLM outputs are lifted into typed compositional derivations -- not claiming that LLMs implement CCG internally, but that their outputs admit a principled, incremental, and auditable CCG reconstruction. Two consequences follow. First, through the Curry-Howard correspondence the lifting extends beyond natural language to the formal languages LLMs also produce -- programming languages such as Solidity, description-logic and query languages such as OWL and SQL -- with the type system varying and the architecture held fixed. Second, the lifting supports two layers of checking: a compositional layer that catches structural failures directly, and a content layer that checks the lifted structure against external knowledge sources, enabling the earliest possible flagging of hallucinated content. The account thereby requires of a producer not cognition but a prefix-driven generative profile. We close with a sketch of synchronous LLM-CCG coupling as one direction the framework opens.

发表机构

  • STAKE Lab, University of Molise(莫利塞大学STAKE实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑