相同公式,不同语义:语言模型是否遵循模态逻辑规范?
Same Formulas, Different Semantics: Do Language Models Follow Modal Logic Specifications?
- Univ. Lille(里尔大学)
- Inria(法国国家信息与自动化研究所)
- CNRS(法国国家科学研究中心)
- Centrale Lille(里尔中央理工学院)
- UMR 9189 - CRIStAL(UMR 9189 - CRIStAL(法国国家科研中心联合研究机构))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究探究语言模型是否遵循模态逻辑规范,构建成对模态问题测试五款模型,发现遵循规定模态语义依赖推理模式与模型身份,发布相关产物。
AI中文摘要:
关于必然性与可能性的推理依赖于对可能世界之间的可达性以及每个世界中存在的对象的假设,因此同一推理可能在一种模态系统中成立,而在另一种模态系统中不成立。在这类问题上评估语言模型,需要测试其判断是否遵循指定的语义而非熟悉的逻辑。我们构建了成对的模态问题,具有相同的前提和猜想,但框架或域条件不同;自动推理验证了相反的标签。平衡核心确保仅语义条件无法单独揭示答案。在该核心上,五款近期模型中有四款在直接提示下的表现低于仅条件基线。然而,启用推理模式后,DeepSeek V4 Flash在未更改的提示下从4.4%提升至88.1%。因此,遵循规定的模态语义强烈依赖于推理模式和模型身份。当省略框架条件时,模型通常达成一致,但最适配不同的熟悉逻辑。我们发布了相关公式、预言机产物、反模型及响应结果。
英文摘要:
Reasoning about necessity and possibility depends on assumptions about accessibility between worlds and about which objects exist at each one. The same inference may therefore hold under one modal system and fail under another. Evaluating language models on such problems requires testing whether their judgments follow the stated semantics rather than a familiar logic. We construct paired modal problems with identical premises and conjecture but different frame or domain conditions; automated reasoning verifies opposite labels. A balanced core prevents the semantic condition alone from revealing the answer. On this core, four of five recent models perform below the condition-only baseline under direct prompting. Yet enabling reasoning mode raises DeepSeek V4 Flash from 4.4% to 88.1% on unchanged prompts. Following stipulated modal semantics thus depends strongly on inference mode as well as model identity. When frame conditions are omitted, models often agree but fit different familiar logics best. We release the formulas, oracle artifacts, countermodels, and responses.