发表机构
East China Normal University; Shanghai Institute of Artificial Intelligence for Education; School of Computer Science and Technology, East China Normal University; Yangzhou University; School of Education and Intelligent Education Research Center, Yangzhou University; School of Data Science and Engineering, East China Normal University(华东师范大学; 上海人工智能教育研究院; 华东师范大学计算机科学与技术学院; 扬州大学; 扬州大学教育学院与智能教育研究中心; 华东师范大学数据科学与工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对复杂多步骤创意任务中LLM作为评判者存在的偏差问题,提出解耦分析与评判环节的自动创意评估器CreaEval,实验显示其性能较次优基线平均提升22.74%且具有泛化性。
AI 中文摘要
将大语言模型(LLM)作为评判者对创意任务进行自动评估仍存在挑战,因为LLM易受冗长偏差、宽松偏差等影响。这类局限性在情境化且流程结构化任务(CGPST)中尤为明显,该任务属于复杂多步骤创意任务,其步骤间依赖关系、高度主观性及宽泛评分范围会导致判断更不稳定、存在更多偏差。现有方法要么依赖特定任务训练,要么直接应用LLM作为评判者,二者均难以在这种复杂性下确保评估可靠性。为填补这些空白,我们提出CreaEval,一种针对CGPST的自动创意评估器,它将典型的LLM作为评判者解耦为分析与评判两个环节。相应地,CreaEval包含两个关键阶段:记忆增强分析阶段,采用SoT-LLM将多步骤响应转换为结构化评估证据,融入跨步骤记忆;基于证据的评判阶段,采用Judge-LLM利用提取的证据进行评判,无需访问原始响应。综合实验表明,在CGPST及两个经典简单创意任务中,CreaEval相较于次优基线的平均性能提升达22.74%,展现出其泛化性。代码可在该https URL获取。
英文摘要
Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to biases such as verbosity bias and leniency bias. Such limitations are particularly evident in Contextually-Grounded and Procedurally-Structured Tasks (CGPST), a complex multi-step creativity task where inter-step dependencies, highly subjectivity, and wide scoring ranges lead to more unstable and biased judgments. Existing approaches either rely on task-specific training or directly apply LLM-as-a-Judge, both of which struggle to ensure reliable evaluation under such complexity. To bridge these gaps, we propose CreaEval, an automated creativity evaluator for CGPST that decouples typical LLM-as-a-Judge into analysis and judging. Correspondingly, CreaEval involves two critical phases: Memory-augmented Analysis, a SoT-LLM converts multi-step responses into structured evaluation evidence, incorporating cross-step memory; and Evidence-based Judging, a Judge-LLM uses the extracted evidence for judging without accessing raw responses. Comprehensive experiments show that CreaEval achieves an average performance improvement of 22.74% over the second-best baselines across CGPST and two classic simple creativity tasks, demonstrating its generalizability. The code is available at https://github.com/Jaong/CreaEval.
CommentsAccepted to EMNLP 2026