发表机构
Haverford College(哈维福德学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究构建含14320个提示的概率算子推理基准,评估29个大语言模型,发现多数模型存在答案偏差,仅9个超随机水平,且各维度均有偏差。
AI 中文摘要
不确定性表达与推理在自然语言中无处不在,对自然语言不确定性表达的有效推理不仅是日常对话的必要部分,也是医疗、法律等高风险领域的核心需求。尽管大语言模型在逻辑推理任务上的评估日益增多,但要从巧妙的表层模式匹配中区分出原则性的符号推理却困难重重。我们推出了一个针对概率算子推理的基准测试,该基准包含14320个程序生成的英文提示,涵盖15种推理模板,系统地改变问题形式、否定策略和表层内容。我们评估了29个模型,发现大多数模型存在与逻辑形式无关的答案偏差,即对“是”或“否”存在系统性偏好,我们将此总结为能力下限:模型在“是”正确项和“否”正确项上准确率的较差值。29个模型中仅有9个超过随机水平。我们还测试了问题形式、动词短语/活动以及提示中使用的姓名性别和来源的变化,发现所有维度都存在偏差。
英文摘要
Both expressions of uncertainty and inferences are ubiquitous in natural language, and valid inferences over natural-language expressions of uncertainty are necessary for not only everyday conversations but also for high-stakes domains such as medicine and law. While large language models are increasingly evaluated on logical reasoning tasks, disentangling principled, symbolic reasoning from clever surface-level pattern matching is fraught with difficulty. We introduce a benchmark for reasoning over probability operators--inference over sentences with gradable epistemic modals (e.g., probably, might, must) containing 14,320 procedurally-generated English prompts across fifteen inference templates, systematically varying question form, negation strategy, and surface content. Evaluating 29 models, we find that most show answer biases independent of the logical form, a systematic preference for Yes or No. We summarize this with a competence floor: the worse of a model's accuracy on Yes-correct and No-correct items. Only 9 of 29 models exceed random chance. We also test variations in question form, verb phrases/activity, and both the gender and origin of names used in the prompts, finding biases across every axis.
CommentsUnder review