arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向现代艺术中的不确定性量化

Toward Uncertainty Quantification in Modern Art

Tirtho Roy, Ushashi Bhattacharjee, Showrav Kumar Saha, Sayantan Chakraborty, Koushik Howlader, Tanusree Bhattacharjee

arXiv 2608.04038首次发表:更新:

AI 中文总结

本研究针对现代艺术动画生成的不确定性问题,提出含多类估计器、分布轮廓等的源盲多种子不确定性识别方案,构建首个语料库,在种子集拓扑分类等任务上表现优异。

AI 中文摘要

当要求文本到视频模型在不同随机种子下对同一现代艺术作品进行动画生成时,模型会返回明显不同的影片,每个种子对应一种解读。由于现代艺术本就具有模糊性,这种差异是有效信号而非噪声。然而,现有的不确定性量化(UQ)方法会将一组生成结果简化为一个离散标量,仅能说明种子间的差异程度,却无法说明差异类型:无法区分是紧凑解读与主导解读加异常值、两种竞争模式,还是弥散型不稳定性,也无法判断该组结果中是否仍包含忠于原作的渲染。我们开展了现代艺术动画生成不确定性结构的首次研究,并提出了一种可复用的源盲多种子不确定性识别方案,该方案包含七个源盲估计器和六个参考感知估计器;一个分布轮廓(包含稳健离散度、异常值影响、显式拓扑、多模态、各向异性、留一种子影响、参考覆盖率);一个分布模型 ablation(涉及vMF、Kent、ACG、Student t、核函数、混合模型);八个识别问题;以及一个艺术品级统计方案。我们构建了首个语料库:由Wan2.1 14B在四个种子下渲染的250条现代艺术作品描述(共1000个视频),涵盖4种编码器,且所用艺术品均未参与生成。作为诊断工具,该方案表现出色:它以0.98的平衡准确率(随机水平为0.25)对种子集拓扑进行分类,在AUROC达1.00时识别出异常值配置,而标量方法仅达0.35,还能从三个种子及跨编码器可靠地将高不确定性艺术品分为参考覆盖型(n=97)和参考缺失型(n=56)两类多样性。

英文摘要

Asked to animate the same modern artwork under different random seeds, a text to video model returns visibly different films, one reading per seed. Because modern art is ambiguous by intent, this disagreement is signal, not noise. Yet prevailing uncertainty quantification (UQ) collapses a set of generations to a dispersion scalar that says how much the seeds differ but not how: it cannot tell a compact interpretation from a dominant reading plus an outlier, two competing modes, or diffuse instability, nor whether the set still contains a rendering faithful to the original. We present the first study of the structure of generative uncertainty for modern art animation, and a reusable protocol for identifying source blind multiseed uncertainty: a suite of seven source blind and six reference aware estimators; a distributional profile (robust spread, outlier influence, explicit topology, multimodality, anisotropy, leave one seed influence, reference coverage); a distribution model ablation (vMF, Kent, ACG, Student t, kernel, mixture); eight identification questions; and an artwork level statistical protocol. We build the first corpus: 250 modern artwork captions rendered by Wan2.1 14B under four seeds (1000 videos) across 4 encoders, artworks withheld from generation. As a diagnostic the protocol succeeds: it classifies seed set topology at balanced accuracy 0.98 (chance 0.25), isolates the outlier configuration at AUROC 1.00 where a scalar reaches only 0.35, and splits high uncertainty artworks into reference covering (n=97) and reference missing (n=56) diversity, reliably from three seeds and across encoders.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑