发表机构
City University of Hong Kong; Northwestern Polytechnical University; The Hong Kong University of Science and Technology; The Hong Kong Polytechnic University(香港城市大学; 西北工业大学; 香港科技大学; 香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Whiteboard基准,交叉评估LLM的想象力与幻觉,发现两者多数子类型呈负相关,挑战了二者同源的观点。
AI 中文摘要
想象力作为大语言模型(LLMs)的一项高级功能,决定了LLM如何创造未见或创新内容的潜力。尽管现有工作已为此能力构建了丰富的创造力基准家族,但它们仅衡量输出偏离常见答案的程度,从不检查这种偏离是否由提示词授权。此外,幻觉作为想象力最邻近的概念,总是在不同的生成结果上通过独立的流程进行衡量,因此,想象力与幻觉源于相同生成机制这一有影响力的论断从未被直接检验。在本文中,我们提出了Whiteboard,这是首个LLM想象力评估基准。其设计遵循了用于衡量人类想象力的权威认知工具:七种基于机制的想象力子类型改编自经典范式,然后与十种支持边界幻觉子类型交叉,并在同一生成结果上联合评分。与以往的创造力或幻觉基准不同,Whiteboard通过显式的支持检查来门控每个想象力分数,并通过可审计的原子矩阵确定性地计算两个轴,主路径上不使用LLM评判器。完整的Whiteboard题库包含1,660个提示;在其共享的80项锚定集上,我们评估了79个最先进的LLM,并针对13,280个人类判断验证了该工具。此外,我们进一步探讨了想象力是否源自与幻觉相同的生成倾向,以及哪些关键因素塑造了它。我们的分析表明幻觉与想象力之间存在反直觉的相关性:大多数子类型耦合为负,且每个锚定项都自身重现了负耦合。
英文摘要
Imagination performs as a high-level function of large language models (LLMs) which determines the potential of how an LLM creates unseen or creative content. While existing works have built a rich family of creativity benchmarks for this ability, they only measure how far an output departs from common answers and never check whether the departure is licensed by the prompt. Moreover, hallucination, the closest neighbor of imagination, is always measured in a separate pipeline on different generations, so the influential claim that imagination and hallucination stem from the same generative mechanism has never been directly testable. In this paper, we propose Whiteboard, the first LLM imagination evaluation benchmark. Its design follows the authoritative cognitive instruments developed to measure human imagination: seven mechanism-grounded imagination subtypes are adapted from classic paradigms, then crossed with ten support-boundary hallucination subtypes and scored jointly on the same generation. Different from previous creativity or hallucination benchmarks, Whiteboard gates every imagination score with an explicit support check and computes both axes deterministically through an auditable atom matrix, with no LLM judge on the primary path. The full Whiteboard item bank contains 1,660 prompts; on its shared 80-item anchor set, we evaluate 79 state-of-the-art LLMs and validate the instrument against 13,280 human judgments. Additionally, we further explore whether imagination derives from the same generative tendency as hallucination and what key factors shape it. Our analysis indicates a counterintuitive correlation between hallucination and imagination: Most of the subtype couplings are negative, every one of the anchor items reproduces the negative coupling on its own.