不是你的本意:LLM 能否遵循指定的否定语义?
Not What You Meant: Can LLMs Follow a Specified Negation Semantics?
- University of Auckland(奥克兰大学)
- Royal Holloway, University of London(伦敦大学皇家霍洛威学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出 NAFBench 基准,评估 LLM 在四种否定语义下的遵循能力,发现开源模型存在明显差距,而前沿模型接近完美,并探讨了求解器委托、微调等缓解方法。
AI中文摘要:
否定在不同领域中并不具有统一的解释。在法律、监管和医学推理中,预期的解释取决于当前生效的解读方式——开放世界与封闭世界、二值与三值、轻信与怀疑。我们研究大型语言模型默认采用哪种否定解读,以及当明确指定不同解读时,它们能否覆盖该偏好。为此,我们引入 NAFBench,一个程序化生成器,用于生成经过求解器验证的实例,涵盖四种语义视角:SLDNF、良基语义(WFS),以及稳定模型语义下的轻信和怀疑推理。该生成器生成具有受控深度、宽度和环结构的地面正规逻辑程序。每个程序在全部四种视角下使用 SWI-Prolog、良基语义求解器和 clingo 求解,产生最多四个不同的标签。随后,程序在多种框架和规则排序下被表述为自然语言,这些框架和排序不改变答案。结果揭示了持续存在的差距。在开源模型中,遵循指定的否定语义仍未解决:最强模型在四种语义视角上得分 59–74%,而最弱模型得分 31–67%。所有模型在超过一半的逻辑等价规则重排上对顺序敏感,而两个较弱的模型经常在良基“未定义”上过度承诺。两个前沿模型在主要固定复杂度评估集上达到 100%,第三个模型 o4-mini 接近完美,仅在良基“未定义”上降至 81%。将推理委托给求解器、在认证轨迹上进行微调,或强制明确的三值判定,均能部分缩小差距。
英文摘要:
Negation does not carry a uniform interpretation across domains. In legal, regulatory, and medical reasoning, the intended interpretation depends on the reading in force -- open- versus closed-world, two- versus three-valued, and credulous versus skeptical. We study which reading of negation large language models adopt by default and whether they can override that preference when a different reading is explicitly specified. To this end, we introduce NAFBench, a procedural generator of solver-certified instances spanning four semantic viewpoints: SLDNF, well-founded semantics (WFS), and credulous and skeptical reasoning under stable-model semantics. The generator emits ground normal logic programs with controlled depth, width, and cycle structure. Each program is solved under all four viewpoints using SWI-Prolog, a well-founded semantics solver, and clingo, yielding up to four divergent labels. The programs are then verbalized into natural language under multiple framings and rule orderings that leave the answer invariant. The results expose a consistent gap. Across open-source models, following a specified negation semantics remains unsolved: the strongest models score 59--74% across the four semantic viewpoints, while the weakest score 31--67%. All models are order-sensitive on more than half of logically identical rule shufflings, while the two weaker models frequently overcommit on well-founded "undefined." Two frontier models reach 100% on the main fixed-complexity evaluation set, and a third, o4-mini, is near-perfect, falling only to 81% on well-founded "undefined." Delegating reasoning to a solver, fine-tuning on certified traces, or forcing an explicit three-valued verdict each partly closes the gap.