ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments
ReliableEval: 通过矩方法进行随机大语言模型评估的配方
机构 * The Hebrew University of Jerusalem(耶路撒冷希伯来大学) ; Google Research(谷歌研究)
AI总结 本文提出ReliableEval方法,通过矩方法评估大语言模型的提示敏感性,发现顶级模型如GPT-4o和Claude-3.7-Sonnet存在显著提示敏感性。
Comments Findings of EMNLP 2025
Journal ref Findings of the Association for Computational Linguistics: EMNLP 2025, pages 11146-11153, Suzhou, China. Association for Computational Linguistics