发表机构
The University of Melbourne; University of New South Wales(墨尔本大学; 新南威尔士大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于LLM评判者的扰动强度测量方法,统一评估输入与思维链扰动下的解释自一致性,发现输入扰动影响更强,且自一致性判断仅在同类扰动内公平。
AI 中文摘要
先前的工作使用表面级扰动方法检验了大语言模型生成解释的自一致性。然而,这些扰动的强度并未被明确测量和控制。在本工作中,我们提出了一种基于大语言模型作为评判者的方法,以统一方式测量输入扰动和思维链扰动的强度。随后,我们在受控强度条件下评估了来自不同大语言模型的解释的自一致性,确保跨扰动类型的公平比较。实验表明,我们提出的基于大语言模型的扰动强度测量方法优于其他基于嵌入和概率的方法,并且输入扰动通常比思维链扰动对大语言模型的影响更强。我们的工作表明,关于模型自一致性的判断仅在相同扰动类型内才是公平的。
英文摘要
Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types. Experiments show that our proposed LLM-based perturbation strength measure outperforms other embedding- and probability-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations. Our work suggests that judgments about a model's self-consistency is fair only within the same perturbation type.
Comments22 pages, 10 figures