Hacking Generative Perplexity: Why Unconditional Text Evaluation Needs Distributional Metrics
破解生成困惑度:为何无条件文本评估需要分布度量
机构 * AITHYRA Institute(AITHYRA研究所)
AI总结 本文指出生成困惑度(gen-PPL)作为非自回归语言模型评估指标存在缺陷,通过构造零参数朴素采样器在LM1B和OpenWebText上达到SOTA gen-PPL但生成不连贯文本,建议采用直接量化生成文本与参考文本分布差异的评估套件。
Comments Presented at the ICML 2026 Workshop on Structured Probabilistic Inference & Generative Modeling (SPIGM), Seoul, South Korea