An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
对LLMs在数学推理中鲁棒性的调查:通过高级数学问题的数学等价转换进行基准测试
机构 * Department of Computer Science University of Illinois Urbana–Champaign(计算机科学系伊利诺伊大学厄巴纳-香槟分校) ; Department of Computer Science Stanford University(计算机科学系斯坦福大学)
AI总结 本文提出了一种新的评估方法,通过数学等价转换的变体测试LLMs的数学推理鲁棒性,发现模型在不同变体上的表现差异显著。
Comments 34 pages, 9 figures