arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06501cs.AIcs.CLcs.MM

多模态大语言模型(MLLMs)能否解码创造性飞跃?推出面向跨概念理解的C4框架

Can MLLMs Decode the Creative Leap? Introducing C4 for Cross-Concept Understanding

Ming Wang, Yuqing Zhang, Tingna Xie, Xiangju Li, Xiaocui Yang, Daling Wang, Shi Feng, Yifei Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究推出受认知启发的C4成语跨概念创造力评估框架,构建含884个案例的C4-Eval集,评估10款MLLMs发现闭源模型准确率最高达50.7%,揭示当前MLLMs创造性解码能力存在显著差距。

中文摘要 AI 辅助

多模态大语言模型(MLLMs)的创造性能力在设计、沟通、教育及人机协作中至关重要,但由于与准确率导向任务相比缺乏明确目标和奖励信号,其创造性能力仍难以评估。跨概念理解是支撑接受性创造力的核心认知能力,它使感知者能从非显而易见但有意义的概念关系中恢复意图含义。我们将项目构建操作化为跨概念编码,将模型推理操作化为跨概念解码。我们推出C4框架,这是一个受认知启发的、用于基于成语(chengyu)的跨概念创造力评估的框架。其编码组件沿人工标注并经第三方审核的跨概念网络中的桥接路径,将目标槽映射为可图像化的替代概念,支持具有明确结构的批量生成,难度由桥接数量和深度索引,且答案明确。利用该框架,我们实例化了C4评估集(C4-Eval),包含184个合成项目及从在线来源收集的37个人工创建的跨概念成语图。我们手动构建并审核了收集到的图的跨概念关系、桥接路径及推理过程。每个C4-Eval项目在5种任务设置中实例化,产生884个主要答案恢复案例。在10个接受评估的MLLMs中,最强的闭源模型达到50.7%的主要准确率,而开源模型的准确率仍显著更低。候选约束能大幅提升准确率,但桥接提示和解释请求仅提供适度提升。这些结果揭示了当前MLLMs在通过跨概念关系解码创造性编码含义方面存在显著差距。代码在补充材料中。

英文摘要

Creative capabilities of MLLMs matter in design, communication, education, and human--AI collaboration, yet remain difficult to evaluate because explicit targets and reward signals are scarce compared with accuracy-oriented tasks. Cross-concept understanding is a core cognitive capacity underlying receptive creativity. It enables a perceiver to recover intended meaning from non-obvious but meaningful conceptual relations. We operationalize item construction as cross-concept encoding and model inference as cross-concept decoding. We introduce C4, a cognition-inspired evaluation framework for Chengyu (Chinese idiom)-based Cross-Concept Creativity. Its encoding component maps target slots to imageable substitute concepts along bridge paths in a manually annotated and third-party-reviewed cross-concept network, enabling batch generation with explicit structure, difficulty indexed by bridge count and depth, and exact answers. Using this framework, we instantiate the C4 Evaluation Set (C4-Eval), comprising 184 synthetic items and 37 human-created cross-concept chengyu figures collected from online sources. We manually construct and review cross-concept relations, bridge paths, and reasoning processes for the collected figures. Each C4-Eval item is instantiated in five task settings, yielding 884 primary answer-recovery cases. Across ten evaluated MLLMs, the strongest closed models reach 50.7% and 48.0% primary accuracy, while open-source models remain substantially lower. Candidate constraints improve accuracy sharply, but bridge hints and explanation requests provide only modest gains. These results expose a substantial gap in how current MLLMs decode creatively encoded meaning through cross-concept relations. The code is in the supplementary material.

发表机构

  • School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
  • School of Computing and Information Systems, Singapore Management University(新加坡管理大学计算与信息系统学院)
  • School of Computer Science and Engineering, Shandong University of Science and Technology(山东科技大学计算机科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

↑