发表机构
Monash University; University College London(莫纳什大学; 伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出工具驱动的符号编辑框架,通过可控修改逻辑算子等结构成分,发现不同LLM在算子编辑下的推理行为不一致,可用于评估模型推理可靠性。
AI 中文摘要
大语言模型(LLMs)的逻辑推理是一项关键能力,它反映了系统运用可靠演绎过程从给定语境中正确推导假设的能力。然而,已有研究表明,LLM的推理对问题表述中微小的表面层面变化十分敏感,这引发了人们对模型是否真正遵循潜在逻辑结构的质疑。研究这种行为颇具挑战性,因为逻辑问题中的符号成分(如算子和谓词)难以在自然语言中进行系统性操控。我们提出了一种工具驱动的框架,用于生成逻辑推理问题的可控、保留标签的编辑。该方法基于一阶逻辑和约束满足问题任务的符号表示运行,能够对逻辑算子及其他结构成分进行针对性修改,之后再将其转换回自然语言。利用该框架,我们在累积和单个算子编辑的条件下评估了多种LLMs,并分析了它们对这些变化的响应行为。我们的定量和定性分析表明,无论模型规模或所属系列如何,LLM在受控算子编辑下的推理行为均不一致:模型有时能正确适应结构变化,但常常无法追踪其逻辑后果。这种自动化压力测试的结果可用于从不同维度评估语言模型,帮助衡量其推理的可靠性。
英文摘要
Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be sensitive to small surface-level variations in problem formulation, raising questions about whether models truly follow the underlying logical structure. Studying this behavior is challenging because the symbolic components of logical problems, such as operators and predicates, are difficult to systematically manipulate in natural language. We introduce a tool-driven framework for generating controlled, label-preserving edits to logical reasoning problems. Our method operates on symbolic representations of first-order logic and constraint satisfaction problem tasks, enabling targeted modifications to logical operators and other structural components before translating them back into natural language. Using this framework, we evaluate various LLMs under cumulative and individual operator edits and analyze their behavior in response to these changes. Our quantitative and qualitative analyses show that LLM reasoning behavior under controlled operator edits is inconsistent, regardless of model size or family: models sometimes adapt correctly to structural changes but often fail to track their logical consequences. The results from this automated stress test enable an evaluation of language models across different dimensions and help measure the reliability of their reasoning.
CommentsAccepted to Findings of EMNLP 2026