arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解耦分析-评判:一种在复杂多步骤创意任务中使用大语言模型(LLM)的自动创意评估器

Decoupled Analysis-Judging: An Automated Creativity Evaluator Using LLMs in Complex Multi-step Creativity Tasks

Xiangyu Wang, Jin Wu, Xiaoyu Li, Chanjin Zheng, Yifeng Zhou

arXiv 2609.03432首次发表:更新:

发表机构

East China Normal University; Shanghai Institute of Artificial Intelligence for Education; School of Computer Science and Technology, East China Normal University; Yangzhou University; School of Education and Intelligent Education Research Center, Yangzhou University; School of Data Science and Engineering, East China Normal University(华东师范大学; 上海人工智能教育研究院; 华东师范大学计算机科学与技术学院; 扬州大学; 扬州大学教育学院与智能教育研究中心; 华东师范大学数据科学与工程学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对复杂多步骤创意任务中LLM作为评判者存在的偏差问题,提出解耦分析与评判环节的自动创意评估器CreaEval,实验显示其性能较次优基线平均提升22.74%且具有泛化性。

AI 中文摘要

将大语言模型(LLM)作为评判者对创意任务进行自动评估仍存在挑战,因为LLM易受冗长偏差、宽松偏差等影响。这类局限性在情境化且流程结构化任务(CGPST)中尤为明显,该任务属于复杂多步骤创意任务,其步骤间依赖关系、高度主观性及宽泛评分范围会导致判断更不稳定、存在更多偏差。现有方法要么依赖特定任务训练,要么直接应用LLM作为评判者,二者均难以在这种复杂性下确保评估可靠性。为填补这些空白,我们提出CreaEval,一种针对CGPST的自动创意评估器,它将典型的LLM作为评判者解耦为分析与评判两个环节。相应地,CreaEval包含两个关键阶段:记忆增强分析阶段,采用SoT-LLM将多步骤响应转换为结构化评估证据,融入跨步骤记忆;基于证据的评判阶段,采用Judge-LLM利用提取的证据进行评判,无需访问原始响应。综合实验表明,在CGPST及两个经典简单创意任务中,CreaEval相较于次优基线的平均性能提升达22.74%,展现出其泛化性。代码可在该https URL获取。

英文摘要

Automated evaluation of creativity tasks remains challenging for LLM-as-a-Judge, as LLM is susceptible to biases such as verbosity bias and leniency bias. Such limitations are particularly evident in Contextually-Grounded and Procedurally-Structured Tasks (CGPST), a complex multi-step creativity task where inter-step dependencies, highly subjectivity, and wide scoring ranges lead to more unstable and biased judgments. Existing approaches either rely on task-specific training or directly apply LLM-as-a-Judge, both of which struggle to ensure reliable evaluation under such complexity. To bridge these gaps, we propose CreaEval, an automated creativity evaluator for CGPST that decouples typical LLM-as-a-Judge into analysis and judging. Correspondingly, CreaEval involves two critical phases: Memory-augmented Analysis, a SoT-LLM converts multi-step responses into structured evaluation evidence, incorporating cross-step memory; and Evidence-based Judging, a Judge-LLM uses the extracted evidence for judging without accessing raw responses. Comprehensive experiments show that CreaEval achieves an average performance improvement of 22.74% over the second-best baselines across CGPST and two classic simple creativity tasks, demonstrating its generalizability. The code is available at https://github.com/Jaong/CreaEval.

CommentsAccepted to EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑