发表机构
Technical University of Munich(慕尼黑工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大语言模型无法组合序数量词的问题,提出可微模糊推理层(DFIL),通过双路径头结合标量瓶颈与隶属函数,实现单调性和t-范数组合推理,无需组合训练数据。
AI 中文摘要
最先进的语言模型被要求解释“大多数学生中的大多数通过了”时,通常回答“大多数”,尽管组合两个“大多数”实例产生的比例更接近“一些”。我们将这一失败归因于架构选择而非数据缺陷:标准分类头将序数类别视为独立标签,没有机制尊重其自然顺序或进行代数组合。我们引入了可微模糊推理层(DFIL),一种双路径预测头,将标准分类器与基于有序隶属函数库的标量瓶颈分支配对。DFIL提供了仅标签头无法继承的两种结构原语:底层数量上的单调性,以及通过t-范数运算进行组合推理而无需任何组合训练数据。标量分支还提供了分析残差的可解释接口。我们在不同的大语言模型家族上对序数自然语言任务实例化了DFIL。
英文摘要
A state-of-the-art language model asked to interpret "most of most students passed" typically answers "most," though composing two instances of "most" yields a proportion closer to "some." We trace this failure to an architectural choice rather than a data deficit: standard classifier heads treat ordinal categories as independent labels, with no mechanism to respect their natural ordering or compose them algebraically. We introduce the Differentiable Fuzzy Inference Layer (DFIL), a dual-path prediction head pairing a standard classifier with a scalar-bottlenecked branch grounded in a bank of ordered membership functions. DFIL supplies two structural primitives that a label-only head cannot inherit: monotonicity in the underlying quantity, and compositional reasoning via t-norm operations without any compositional training data. The scalar branch additionally provides an interpretable interface for analyzing residual errors. We instantiate DFIL on ordinal natural-language tasks across diverse LLM families.