发表机构
University of California, Los Angeles(加利福尼亚大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究发现进行体悖论相关的基准失效,构建词汇匹配最小对并评估多步推理,揭示模型的充分性偏差等失效模式,表明部分大模型体分类性能可媲美人类标注者。
AI 中文摘要
进行体悖论为组合语义分析提供了有用的测试。近期研究构建了一个自然语言推理(NLI)基准,报告称模型常从进行体描述中推断出已完成的完结事件,并将此行为归因于目的论偏差;该研究还指出提示干预会引发校准危机。我们重新审视该基准及相关结论,发现其受概念与评估规格错误的显著影响,共识别出三类概念规格错误,其中体简化错误影响了基准构建、分析、实验及结论。按照严格的NLI标准,A组实例中有76%未明确排除完结性;经母语标注,A组示例的38%、C组示例的29%被判定为允许其他解读。为控制这些问题及词汇变异,我们构建了词汇匹配的最小对;在评估层面,我们将事件语义NLI定义为多步推理问题,同时评估中间语义决策与最终预测。结果显示,模型常不确认完结性却接受对应的简单过去时假设,此模式被我们表征为充分性偏差;我们还发现提示干预会引发标签间的决策偏移,却无法可靠提升底层语义理解与推理能力。中间分析与oracle引导分析识别出另外两类失效模式:组合体分类错误、对与表面相关答案的表面形式吸引。我们对Qwen-7B(采用合适提示)、GPT-5.4及Qwen-72B的实验为体分类的语境敏感性提供了初步证据,表明这些模型可达到与人类标注者相当的性能。
英文摘要
The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. It further argues that prompting interventions cause a Calibration Crisis. We reexamine the benchmark and conclusions and show that it is substantially affected by conceptual and evaluation mis-specifications. We identify three conceptual mis-specifications. In particular, Aspectual Reduction affects the benchmark construction, analysis, experiments, and conclusions. Under a strict NLI standard, 76% of Group A instances do not explicitly rule out culmination. In our native-speaker annotation, 38% of Group A examples and 29% of the Group C examples were judged to permit an alternative interpretation. To control these issues and lexical variation, we construct Lexically Matched Minimal Pairs. At the evaluation level, we formulate event-semantic NLI as a Multi-step Reasoning Problem and assess both intermediate semantic decisions and final predictions. Our results show that models often do not affirm culmination but nevertheless accept the corresponding simple-past hypothesis, a pattern we characterize as Sufficiency Bias. We further show that prompting interventions produce a Decision Shift among labels without reliably improving the underlying semantic understanding and reasoning. Intermediate and oracle-guided analyses identify two additional failure modes: errors in compositional aspectual classification and Surface-form Attraction toward surface-associated answers. Our experiments on Qwen-7B with suitable prompts, GPT-5.4, and Qwen-72B provide initial evidence for the context sensitivity of aspectual classification and suggest that these models can achieve performance comparable to that of human annotators.