发表机构
Tehran Institute for Advanced Studies, Khatam University; Cardiff University(哈塔姆大学德黑兰高等研究院; 卡迪夫大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM评估波斯语短篇文学文本创造性的维度依赖问题,提出CLIN框架,用三类代理指标评估TTCT衍生的创造性维度,实现了与零样本LLM相当或更优的人类对齐度且评估成本更低。
AI 中文摘要
评估大语言模型(LLM)输出的创造性仍具挑战性,因为创造性是多维度且以人类为中心的。我们研究LLM在多种评估策略和提示表述下对低资源语言波斯语短篇文学文本的评估可靠性。发现LLM与人类的一致性在不同维度差异显著:在结构化的托兰斯创造性思维测验(TTCT)衍生属性(如原创性、流畅性、精细性)上的对齐度更强,但在更主观的维度(尤其是情感和吸引力)上的对齐度弱得多。判断也对提示表述敏感,而少样本提示、集成和多智能体辩论未带来一致的改进。受这种维度依赖行为的启发,我们研究结构化创造性维度是否可通过简单、可解释的代理指标近似。我们提出CLIN,它使用感知主题的新颖性评估原创性,使用上下文词汇聚类评估流畅性,使用词汇多样性评估精细性,分别评估三个TTCT衍生维度。这些代理指标在我们的设置中达到与最强零样本LLM评估器相当或更好的人类对齐度,同时需要低得多的评估成本。
英文摘要
Evaluating creativity in large language model (LLM) outputs remains challenging because creativity is multidimensional and human-centered. We examine how reliably LLMs evaluate short literary text in Persian, a low-resource language, across multiple evaluation strategies and prompt formulations. We find that LLM-human agreement varies substantially across dimensions: alignment is stronger for structured TTCT-derived properties such as Originality, Fluency, and Elaboration, but considerably weaker for more subjective dimensions, particularly Emotion and Attractiveness. Judgments are also sensitive to prompt formulation, while few-shot prompting, ensembling, and multi-agent debate provide no consistent improvement. Motivated by this dimension-dependent behavior, we investigate whether structured creativity dimensions can instead be approximated using simple, interpretable proxy metrics. We introduce CLIN, which evaluates three TTCT-derived dimensions separately using topic-aware novelty for Originality, contextual lexical clustering for Fluency, and lexical diversity for Elaboration. These proxies achieve human alignment comparable to or better than the strongest zero-shot LLM judge in our setting while requiring substantially lower evaluation cost.
CommentsAccepted to Findings of EMNLP 2026