arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可微模糊推理层:用于大语言模型的单调、组合序数推理头

Differentiable Fuzzy Inference Layer: A Monotone, Compositional Ordinal Reasoning Head for Large Language Models

Zhen Zhang, Amr Alanwar

arXiv 2609.26113首次发表:更新:

发表机构

Technical University of Munich(慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型无法组合序数量词的问题,提出可微模糊推理层(DFIL),通过双路径头结合标量瓶颈与隶属函数,实现单调性和t-范数组合推理,无需组合训练数据。

AI 中文摘要

最先进的语言模型被要求解释“大多数学生中的大多数通过了”时,通常回答“大多数”,尽管组合两个“大多数”实例产生的比例更接近“一些”。我们将这一失败归因于架构选择而非数据缺陷:标准分类头将序数类别视为独立标签,没有机制尊重其自然顺序或进行代数组合。我们引入了可微模糊推理层(DFIL),一种双路径预测头,将标准分类器与基于有序隶属函数库的标量瓶颈分支配对。DFIL提供了仅标签头无法继承的两种结构原语:底层数量上的单调性,以及通过t-范数运算进行组合推理而无需任何组合训练数据。标量分支还提供了分析残差的可解释接口。我们在不同的大语言模型家族上对序数自然语言任务实例化了DFIL。

英文摘要

A state-of-the-art language model asked to interpret "most of most students passed" typically answers "most," though composing two instances of "most" yields a proportion closer to "some." We trace this failure to an architectural choice rather than a data deficit: standard classifier heads treat ordinal categories as independent labels, with no mechanism to respect their natural ordering or compose them algebraically. We introduce the Differentiable Fuzzy Inference Layer (DFIL), a dual-path prediction head pairing a standard classifier with a scalar-bottlenecked branch grounded in a bank of ordered membership functions. DFIL supplies two structural primitives that a label-only head cannot inherit: monotonicity in the underlying quantity, and compositional reasoning via t-norm operations without any compositional training data. The scalar branch additionally provides an interpretable interface for analyzing residual errors. We instantiate DFIL on ordinal natural-language tasks across diverse LLM families.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑