从词元概率到语义约束:迈向语言模型的声明式概率评估
From Token Probabilities to Semantic Constraints: Towards Declarative Probabilistic Evaluation of Language Models
浏览论文内容
中文总结 AI 辅助
本文提出ModelLog,一个声明式概率评估框架,通过符号约束衡量模型分布满足程度,发现否定、互斥和一致性任务中的系统性失败,并揭示评估分数可作为损失连接评估与学习。
中文摘要 AI 辅助
尽管大型语言模型已经取得了快速进步,但关于如何评估它们所获得的知识和推理能力,以及这些评估如何与预训练中使用的学习信号相关联,仍然存在许多基本问题。在本文中,我们提出了ModelLog,一个用于预训练评估的声明式概率框架,该框架使模型行为的语义结构变得明确,并为将评估与学习联系起来提供了新的形式化工具。ModelLog将评估目标指定为对词元级别预测的符号约束,并衡量模型分布满足这些约束的强度。我们通过一套新的任务来探索该框架,这些任务针对否定、互斥性和一致性,发现了难以仅通过词元似然或答案准确率来表征的系统性失败。我们进一步表明,这些评估分数也可以解释为损失,其梯度反映了逻辑强度、信息量和变量级别的敏感性。这通过共享语义将评估和学习联系起来,提出了能够诊断模型行为同时也有助于阐明学习语义结构的评估方法。
英文摘要
While Large Language Models have improved rapidly, many fundamental questions remain about how to evaluate the knowledge and reasoning abilities they acquire, and how such evaluations relate to the learning signals used in pre-training. In this paper, we propose ModelLog, a declarative probabilistic framework for pre-training evaluation that makes the semantic structure of model behavior explicit and provides new formal tools for relating evaluation to learning. ModelLog specifies evaluation targets as symbolic constraints over token-level predictions and measures how strongly a model's distribution satisfies those constraints. We explore the framework through a new suite of tasks targeting negation, mutual exclusivity, and consistency, finding systematic failures that are difficult to characterize through token likelihood or answer accuracy alone. We further show that these evaluation scores can also be interpreted as losses, whose gradients reflect logical strength, informativeness, and variable-level sensitivity. This links evaluation and learning through a shared semantics, suggesting evaluation methods that diagnose model behavior while also helping to clarify the semantic structure of learning.
发表机构
- Allen Institute for AI(艾伦人工智能研究所)
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。