arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过聚类截断提高自回归文本到图像生成中的样本多样性

Improving Sample Diversity in Autoregressive Text-to-Image Generation via Cluster Truncation

Trang Nguyen, Shuang Wu, Runyan Tan, Phillip Howard

arXiv 2607.10535首次发表:更新:

发表机构

University of Massachusetts Amherst; Thoughtworks(马萨诸塞大学阿默斯特分校; Thoughtworks公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究自回归图像生成模型中样本多样性问题,指出其关键属性限制现有方法。提出无p聚类新解码策略,在四个自回归T2I模型和两个数据集上评估,该策略能解锁最大多样性并保持图像质量与提示对齐。

AI 中文摘要

虽然扩散模型在文本到图像(T2I)生成中实现了最先进的图像质量,但最近的工作表明它们存在样本多样性崩溃的问题。在这项工作中,我们研究自回归(AR)图像生成模型是否能推动图像质量和样本多样性之间的帕累托前沿。随着质量和效率的最新进展,AR模型已成为基于扩散的图像生成的可行替代方案。受AR范式中更好的多样性-质量权衡潜力的推动,我们对AR图像生成模型中的样本多样性进行了首次系统研究。我们表明,AR图像生成的两个关键属性,即持续高的令牌级熵和视觉令牌空间中的大量冗余,限制了现有令牌级解码方法增强多样性的有效性。因此,我们提出了无p聚类,这是一种新的解码策略,它在聚类级别而不是令牌级别执行基于熵的截断采样。我们使用一套全面的指标,包括图像质量、提示对齐和多样性,在四个自回归T2I模型和两个数据集上评估了我们的方法和基线解码方法。我们的结果表明,无p聚类在大多数评估的自回归T2I模型和数据集上解锁了最大的多样性,同时保持了图像质量和提示对齐。

英文摘要

While diffusion models achieve state-of-the-art image quality for text-to-image (T2I) generation, recent work has demonstrated that they suffer from sample diversity collapse. In this work, we investigate whether autoregressive (AR) image generation models can push the Pareto frontier between image quality and sample diversity. With recent advances in quality and efficiency, AR models have emerged as a viable alternative to diffusion-based image generation. Beyond enabling new use cases such as interleaved image-text generation, their sequential generation process makes them compatible with a wide range of token-based decoding strategies originally developed to improve diversity in text generation. Motivated by the potential of a better diversity-quality tradeoff in the AR paradigm, we present the first systematic study of sample diversity in AR image generation models. We show that two key properties of AR image generation, persistently high token-level entropy and substantial redundancy in visual token spaces, limit the effectiveness of existing token-level decoding methods for diversity enhancement. We therefore propose $p$-less cluster, a new decoding strategy that performs entropy-based truncation sampling at cluster level rather than at token level. We evaluate our approach and baseline decoding methods across four autoregressive T2I models and two datasets using a comprehensive suite of metrics spanning image quality, prompt alignment, and diversity. Our results show that $p$-less cluster unlocks the greatest diversity across most evaluated autoregressive T2I models and datasets while maintaining image quality and prompt alignment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑