arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22554cs.AIcs.CLcs.LG

相同问题,不同答案:超越准确率评估大语言模型的可靠性

Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy

Kazem Faghih, Yize Cheng, Shoumik Saha, Mobina Pournemat, Armin Gerami, Soheil Feizi

首次发表
浏览论文内容

中文总结 AI 辅助

研究大语言模型在不同等效表述问题下的答案变化,发现其输出依赖措辞,准确率指标掩盖不稳定性,实例级行为不稳定,提出自释义策略可恢复潜在知识提升性能,强调评估一致性对展现模型可靠性的重要性。

中文摘要 AI 辅助

大语言模型在基准测试中常取得高准确率,但对于相同问题的不同等效表述,其应用知识的可靠性仍不明晰。本文研究了在事实问答和数学推理任务中,模型答案在保持意义的释义下如何变化。跨越四个基准测试和13个模型,发现模型输出常依赖提示的确切措辞。整体准确率变化不大,但实例级行为不稳定,许多问题的答案因措辞而异,不匹配率超23%。以原始形式正确回答的问题为例,答案翻转率显示出更大问题。同时发现模型常能为问题的至少一种释义给出正确答案。基于此,一种简单的自释义策略可部分恢复潜在知识并提升推理性能。这些发现表明标准准确率指标可能掩盖不稳定性,评估等效输入的一致性能更清晰地展现大语言模型的可靠性。

英文摘要

Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change under meaning-preserving paraphrases across factual question answering and mathematical reasoning tasks. Across four benchmarks and 13 models, we find that model outputs frequently depend on the exact wording of the prompt. While overall accuracy typically changes only modestly across paraphrases, instance-level behavior is far less stable: for many questions, models alternate between correct and incorrect answers depending on phrasing, with mismatch rates reaching more than 23%. Conditioning on questions that are answered correctly in their original form reveals even larger failures measured by answer flip rates, showing that single-prompt correctness is often a poor indicator of reliability. At the same time, we find that models often produce a correct answer for at least one paraphrase of a question, suggesting that the underlying knowledge is present but inconsistently retrieved. Building on this observation, we show that a simple self-paraphrasing strategy can partially recover this latent knowledge and improve performance at inference time. Together, these findings suggest that standard accuracy metrics can mask substantial instability, and that evaluating consistency across equivalent inputs provides a clearer picture of LLM reliability.

发表机构

  • University of Maryland, College Park(马里兰大学帕克分校)

机构由 AI 辅助整理,请以论文原文为准。

↑