用多样性轮廓评估AI生成内容的多样性
Evaluating the Diversity of AI-Generated Content with Diversity Profiles
- Tsinghua University(清华大学)
- MBZUAI(穆罕默德·本·扎耶德人工智能大学)
- Microsoft Research(微软研究院)
- University of Cambridge(剑桥大学)
- McGill University(麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对现有AI生成内容多样性评估标量指标的矛盾性与局限性,提出曲线值的多样性轮廓框架,为生成式AI评估提供更透明、感知分辨率的多样性比较方案。
AI中文摘要:
多样性是评估生成式人工智能(AI)系统的基本标准,但其测量本质上仍存在模糊性。现有方法通常将生成样本表示在嵌入空间中,计算成对距离或相似度,并将其聚合为单个标量分数。这种标量总结虽方便,但往往编码不同的归纳偏差,可能对相同样本集产生矛盾的排名。本文认为,将AI生成内容的多样性评估简化为单个数字本质上是不充分的。我们首先回顾代表性的多样性指标,然后从两个互补角度诊断其局限性:公理分析显示,没有代表性的标量指标能同时满足所有理想属性;实证分析显示,高维表示空间可诱导出集中的、依赖模态的距离分布。为解决这些问题,我们提出多样性轮廓:曲线值的、感知条件的总结,在指定的表示以及距离或核函数下,评估参数化多样性族在一系列阈值、尺度、指数或阶数上的情况。多样性轮廓可揭示比较是否在不同分辨率下稳健,还是取决于任意参数选择。我们为几个代表性度量族实例化轮廓,并展示其在生成式AI评估中的实际应用。总体而言,多样性轮廓为比较AI生成内容的多样性提供了更透明、感知分辨率的框架。
英文摘要:
Diversity is a fundamental criterion for evaluating generative artificial intelligence (AI) systems, yet its measurement remains inherently ambiguous. Existing approaches typically represent generated samples in an embedding space, compute pairwise distances or similarities, and aggregate them into a single scalar score. Such scalar summaries are convenient, but they often encode different inductive biases and may yield contradictory rankings of the same sample sets. In this paper, we argue that diversity evaluation for AI-generated content is intrinsically under-specified when reduced to a single number. We first review representative diversity metrics, and then diagnose their limitations from two complementary perspectives: an axiomatic analysis showing that no representative scalar metric satisfies all desirable properties simultaneously, and an empirical analysis showing that high-dimensional representation spaces can induce concentrated, modality-dependent distance distributions. To address these issues, we propose diversity profiles: curve-valued, condition-aware summaries that evaluate a parameterized diversity family across a range of thresholds, scales, exponents, or orders under a specified representation and distance or kernel function. Diversity profiles reveal whether a comparison is robust across resolutions or instead depends on an arbitrary parameter choice. We instantiate profiles for several representative metric families and demonstrate their practical use in generative AI evaluation. Overall, diversity profiles provide a more transparent and resolution-aware framework for comparing the diversity of AI-generated content.