发表机构
New York University; Center for Data Science, NYU Shanghai; Georgia Institute of Technology(纽约大学; 上海纽约大学数据科学中心; 佐治亚理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对开放文本生成评估忽略全局连贯性的问题,提出连贯性感知的分布指标CHORD,利用冻结LLM隐状态和RBF-MMD检测多种连贯性失败,实验证明其优于现有基线并与人类判断高度一致。
AI 中文摘要
现有的开放文本生成评估指标在通用表示空间中衡量似然度、词汇多样性或分布相似性,但它们可能遗漏质量的基本维度。一个突出的盲点是全局连贯性:生成的段落可能在局部流畅,但在全局上矛盾、因果不一致或主题脱节。此类失败仍能保留现有指标所依赖的标记级和词汇统计。我们识别出表示为检测这些失败的核心瓶颈,并引入CHORD(连贯性感知的隐状态开放生成参考距离),一种连贯性敏感的分布指标。CHORD使用连贯性引导提示在冻结LLM的隐状态空间中对生成语料和人类写作语料进行编码,并使用带RBF核的MMD比较所得分布。为验证该指标响应连贯性退化而非通用文本变化,我们构建了一个反事实评估套件,将分级连贯性退化扰动与保持意义的控制配对。CHORD选择性地检测关系、话语、结构和混合失败,而这些失败是困惑度、熵、MAUVE、FBD和基于MMD的基线要么遗漏,要么无法与良性改写区分。因子消融表明表示是连贯性敏感性的主要来源,而RBF-MMD在相关区分变得可见后提高了样本效率。更大的骨干模型捕捉更细粒度的区分,但连贯性提示仅在骨干模型能遵循该提示时提高选择性。在无条件生成和前缀延续中,CHORD产生的模型排名与人类对输出是否有意义且看起来像人类写作的判断强烈一致。总之,这些结果确立了表示设计作为可靠分布评估的核心。
英文摘要
Existing open-ended generation metrics measure likelihood, lexical diversity, or distributional similarity in generic representation space, yet can miss fundamental dimensions of quality. A prominent blind spot is global coherence: a generated passage may be locally fluent while remaining globally contradictory, causally inconsistent, or topically disconnected. We identify representation as a central bottleneck in detecting these failures and introduce CHORD (Coherence-aware Hidden-state Open-generation Reference Distance), a coherence-sensitive distributional metric. CHORD encodes generated and human-written corpora in the hidden-state space of a frozen LLM using a coherence-eliciting prompt, and compares the resulting distributions using RBF-MMD. To test coherence sensitivity and selectivity, we construct a counterfactual evaluation suite pairing graded coherence-degrading perturbations with meaning-preserving controls. CHORD selectively detects relation, discourse, structural, and mixture failures that perplexity, entropy, MAUVE, FBD, and MMD-based baselines either miss or cannot separate from benign rewriting. Factorial ablations show that representation is the primary source of coherence sensitivity, while RBF-MMD improves sample efficiency. Larger backbones capture finer-grained distinctions, but coherence prompting improves selectivity only when the backbone can follow the prompt. On unconditional generation and prefix continuation, CHORD yields model rankings that strongly align with human judgments of whether outputs make sense and appear human-written. Together, these results establish representation design as central to reliable distributional evaluation. Code: https://github.com/MAPS-research/CHORD. Experiments: https://github.com/MAPS-research/CHORD-Experiment.
CommentsPreprint. 41 pages, 13 figures