发表机构
Santa Clara University(圣克拉拉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究测试了14个语言模型在成对力学问题上的表现,发现模型常能识别问题缺陷但仍报告为已解决,因此评估需同时考察求解与拒绝能力。
AI 中文摘要
语言模型起草工程计算,但答案的准确性并不能表明它们是否拒绝一个不可能的问题。我们测试了14个模型在30对力学问题上的表现,每对包含一个有效版本和一个通过改变给定值或假设而变得不可能的版本。两个独立的求解器验证了每个答案要点,并表明每个有缺陷的问题在物理上都是不可能的。我们分别对有效问题的求解和对其有缺陷对应问题的拒绝进行了评分。每个回复都需要一个“已解决”或“无法解决”的状态;拒绝意味着“无法解决”或保留答案。初始提示并未警告问题可能是有缺陷的。在最近的三个模型中,90个回复中有12个未能拒绝一个有缺陷的问题。在这些回复中,有11个模型陈述了缺陷,回答了一个修正后的问题,但仍然将原始问题报告为“已解决”,根据人工智能评分者和数值检查。我们后来重新测试了来自一个提供商的四个模型,提供“有缺陷”而不是“无法解决”的选项,并要求它们指出并解释缺陷。三个模型在拒绝率上显示出统计显著的提高,但有效问题的求解率在三个模型中有所下降。因此,评估需要同时对两个版本进行评分,并区分缺陷识别与报告状态。
英文摘要
Language models draft engineering calculations, but answer accuracy does not show whether they reject an impossible problem. We tested 14 models on 30 pairs of mechanics problems, each with a valid version and one made impossible by changing a given value or assumption. Two independent solvers verified every answer key and showed that each flawed problem was physically impossible. We scored solving of valid problems separately from rejection of their flawed counterparts. Each reply required a "solved" or "cannot solve" status; rejection meant "cannot solve" or withholding an answer. The initial prompts did not warn that problems could be flawed. Across three recent models, 12 of 90 replies failed to reject a flawed problem. In 11 of these replies, the model stated the flaw, answered a corrected problem and still reported the original as "solved", according to artificial intelligence raters and numerical checks. We later retested four models from one provider, offering "flawed" instead of "cannot solve" and asking them to name and explain the defect. Three models showed statistically significant increases in rejection, but valid-problem solving fell in three. Evaluations therefore need to score both versions and distinguish flaw recognition from the reported status.
Comments36 pages, 6 figures, 3 tables; Supplementary Information included as an appendix; figure source data as ancillary files