发表机构
Ruhr University Bochum; University of Duisburg-Essen(波鸿鲁尔大学; 杜伊斯堡-埃森大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大型语言模型(LLMs)的评估问题,借鉴哲学认识论概念提出认知多样性评估维度,构建初步框架并发现前沿LLMs存在认知狭隘性,建议评估需纳入该维度。
AI 中文摘要
大型语言模型(LLMs)越来越多地被用于不仅检索信息,还用于回答问题、进行解释、教学及支持探究。在这类场景中,评估不能仅靠准确率或对齐性来完成。一个系统可能给出正确答案,却仍限制用户获取其他有效答案、解释或推理路径的途径。借鉴哲学与社会认识论中更宽泛的认知多样性概念,我们在LLMs语境中将其形式化为LLMs向用户呈现的有效答案、解释及推理路径的范围。我们认为,在LLMs用于支持知识密集型任务的场景中,认知多样性是一个有用的评估维度。我们提出了一个用于概念化和测量LLMs中认知多样性的初步框架,并在两个领域中对其进行了操作化。我们发现,前沿LLMs通常表现出认知狭隘性,反复将庞大的有效答案空间压缩到小型的规范子集上。这些发现表明,LLM评估应超越以准确率为导向的范式,将认知多样性视为模型能力的一个重要维度。
英文摘要
Large language models (LLMs) are increasingly used not only to retrieve information, but to answer questions, explain, teach, and support inquiry. In such settings, evaluation cannot be exhausted by accuracy or alignment alone. A system may give a correct answer while still narrowing users' access %to knowledge. to alternative valid answers, explanations, or reasoning routes. Drawing on the broader notion of epistemic diversity in philosophy and social epistemology, we formalize it in the context of LLMs as the range of valid answers, explanations, and reasoning routes that an LLM exposes to users. We argue that epistemic diversity is a useful evaluation dimension for settings where LLMs are used to support knowledge-intensive tasks. We propose a preliminary framework for conceptualizing and measuring epistemic diversity in LLMs, and operationalize it in two domains. We find that frontier LLMs often exhibit epistemic narrowness, repeatedly collapsing large valid answer spaces onto small canonical subsets. These findings suggest that LLM evaluation should move beyond accuracy-oriented paradigms and treat epistemic diversity as an important dimension of model capability.