发表机构
LMU Munich; Munich Center for Machine Learning (MCML)(慕尼黑大学; 慕尼黑机器学习中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于参考的框架,从三个视角分析开放式文本生成的连贯性与多样性,实验表明多样性对齐和均值比较能捕捉质量变异,参考似然性与评分正相关,为评估生成质量提供结构化方法。
AI 中文摘要
评估开放式文本生成涉及理解续写的不同属性与其感知质量之间的关系。我们提出了一个基于参考的框架,通过三个视角来考察连贯性和多样性:将其演变与人类轨迹对齐,将其摘要与同一提示的人类续写进行比较,以及估计其在人类参考分布下的似然性。使用人类质量评分的实验表明,基于多样性的对齐和基于均值的比较能够捕捉与质量相关的变异,尽管这些比较并未确立时间对齐相对于更简单基线的预测优势。参考似然性也与评分呈正相关,结果因参考配置和评分范围而异。综合来看,这些分析提供了一种结构化的方式来考察测量的连贯性和多样性与人类判断的关系,同时将相似性与人类参考本身与质量区分开来。代码和分析资源可在 https://github.com/EstebanGarces/likely_human 获取。
英文摘要
Evaluating open-ended text generation involves understanding how different properties of a continuation relate to its perceived quality. We present a reference-based framework for examining coherence and diversity through three perspectives: aligning their evolution with human trajectories, comparing their summaries with a human continuation of the same prompt, and estimating their likelihood under a human reference distribution. Experiments with human quality ratings suggest that diversity-based alignment and mean-based comparisons capture quality-related variation, although the comparisons do not establish a predictive advantage for temporal alignment over simpler baselines. Reference likelihood also shows positive associations with ratings, with results varying across reference configurations and scoring horizons. Together, these analyses provide a structured way to examine how measured coherence and diversity relate to human judgments, while distinguishing similarity to human references from quality itself. Code and analysis resources are available at https://github.com/EstebanGarces/likely_human.
CommentsAccepted at INLG 2026