发表机构
Carnegie Mellon University; The Hebrew University of Jerusalem(卡内基梅隆大学; 耶路撒冷希伯来大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出StructSim,通过将想法分解为目的、机制和实现组件并构建多层概念图,以结构表示度量想法相似性,较嵌入基线提升31%的专家判断一致性,支持大规模想法多样性评估。
AI 中文摘要
度量想法相似性对于创造力评估至关重要,尤其是在大型语言模型(LLM)使得想法生成规模日益增大的背景下。然而,文本嵌入将想法压缩为单一向量,难以捕捉结构相似性,包括核心组件与支持组件之间的部分重叠以及不同抽象层次上的差异。我们提出了一种共享的结构表示方法,将想法分解为目的、机制和实现组件,并在多层概念图中组织相关组件。基于该表示,我们定义了成对相似性和集合级机制覆盖度的度量,以评估想法多样性。我们使用受控想法三元组和12位专家的评估来验证我们的方法,重点关注核心机制、实现和支持组件上的差异。与嵌入基线相比,我们的方法在结构相似性上与专家判断的一致性提高了31%,并能更好地反映专家对想法集合覆盖度的评估,从而支持想法相似性和多样性的可扩展评估。
英文摘要
Measuring idea similarity is fundamental to creativity evaluation, especially as LLMs enable idea generation at increasing scale. However, text embeddings collapse an idea into a single vector, making it difficult to capture structural similarity, including partial overlap across core and supporting components and differences across levels of abstraction. We introduce a shared structural representation that decomposes ideas into purpose, mechanism, and implementation components and organizes related components in a multi-layer concept graph. From this representation, we define measures of pairwise similarity and set-level mechanism coverage for assessing idea diversity. We evaluate our approach using controlled idea triples and assessments from 12 experts, focusing on differences in core mechanisms, implementations, and supporting components. Our method improves alignment with expert judgments of structural similarity by 31% over the embedding baseline and better reflects expert assessments of idea set coverage, supporting scalable evaluation of idea similarity and diversity.