一种受认知启发的用于评估隐喻解释的多维框架
A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations
浏览论文内容
中文总结 AI 辅助
该研究提出受认知启发的多维框架评估隐喻解释,经标注研究验证其多维性等特性,自动评估流程可部分恢复该结构,表明多维评估更具诊断价值,自动评估器需以保留人类判断结构为评判标准。
中文摘要 AI 辅助
当前对隐喻解释的评估主要依赖整体质量评分,几乎无法揭示解释质量的构成结构,也无法体现人类判断的一致与分歧之处。我们提出一种受认知启发的框架,将隐喻解释质量分解为六个基于理论的维度。在一项包含11200次评分的密集标注研究中,我们发现:(i)解释质量确实具有多维性;(ii)标注者的分歧是系统性的而非随机的;(iii)这六个维度可归为一个共享集群和两个独立的判断轴。一项探索性可行性研究进一步表明,标准自动评估流程可恢复该结构的部分内容,能较好预测最具区分度的维度,且其误差与人类判断的分歧相关。综上,这些结果表明,多维评估比整体评分能提供更丰富的诊断见解,对于开放式生成任务的自动评估器,应根据其保留人类判断结构的程度进行评判。
英文摘要
Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge. We introduce a cognitively motivated framework that decomposes metaphor explanation quality into six theoretically grounded dimensions. In a dense annotation study (11,200 ratings), we find that: {\bfseries(i)} explanation quality is genuinely multidimensional; {\bfseries(ii)} annotator disagreement is systematic rather than random; and {\bfseries(iii)} the six dimensions collapse into a shared cluster and two independent axes of judgment. An exploratory feasibility study further shows that a standard automatic evaluation pipeline can recover parts of this structure, predicting the most discriminative dimensions well while its errors correlate human (dis)agreement. Together, these results suggest that multidimensional evaluation offers richer diagnostic insight than holistic ratings, and that automatic evaluators for open-ended generation tasks should be judged on how well they preserve the structure of human judgment.
发表机构
- Constructor University(康斯特大学)
- University of Turin(都灵大学)
- Örebro University(厄勒布鲁大学)
- University of Salerno(萨勒诺大学)
机构由 AI 辅助整理,请以论文原文为准。