arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30448cs.PF

存储还是重新生成?大规模AI生成内容的成本模型

To Store or To Regenerate? A Cost Model for AI-Generated Content at Scale

  • Harvard University(哈佛大学)
  • University of Virginia(弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Yunjia Zheng, Zirui Wang, Haoran Ni, Tingfeng Lan, Zhaoyuan Su, Yue Cheng, Juncheng Yang

AI总结:

本文提出成本模型,比较AI生成内容的持久存储与按需重新生成,发现基于提示的重新生成在2040年前不具成本优势,而缓存结合中间表示(IR)的重新生成可降低成本至少2倍,并在大规模生产轨迹中验证了其有效性。

AI中文摘要:

AI生成内容正成为快速增长的数字制品类别。由于这些制品随时间不断累积,其指数级增长给运营商和社会带来了巨大的存储、能源和基础设施成本。与此同时,GPU计算成本随着每一代硬件更新而持续快速下降。这种分化引发了一个基本问题:何时按需重新生成比持久存储更便宜?本文开发了一个成本模型,用于比较AI生成制品的持久存储与按需重新生成。该模型考虑了语料库增长、HDD和磁带价格趋势、驱动器更换、电力、请求偏斜、缓存、生成器FLOPs以及未来GPU性价比改进。对于图像生成,我们的分析表明,基于提示的重新生成直到大约2040年才会变得比存储更便宜,因为每次缓存未命中仍必须重新运行完整的提示到制品生成流程。我们观察到,广泛使用的基于扩散的生成模型在潜在空间中操作,这为成本权衡创造了一个替代点:运营商不是存储最终制品或仅存储提示,而是可以存储紧凑的中间表示(IR)并执行廉价的按需解码。我们的分析表明,缓存结合基于IR的重新生成显著降低了存储和计算成本,即使今天也比全对象存储和基于提示的重新生成至少便宜2倍。在具有20.7亿请求的生产图像轨迹上,同样的结论成立:基于提示的重新生成比存储贵超过100倍,而基于IR的重新生成将总成本降低到约为全对象存储的一半,同时保持交互式未命中延迟。

英文摘要:

AI-generated content is becoming a rapidly growing class of digital artifacts. Because these artifacts accumulate over time, their exponential growth creates a substantial storage, energy, and infrastructure cost for operators and society. At the same time, GPU compute cost continues to fall rapidly with each hardware generation. This divergence raises a fundamental question: when does on-demand regeneration become cheaper than persistent storage? This paper develops a cost model for comparing persistent storage and on-demand regeneration for AI-generated artifacts. The model accounts for corpus growth, HDD and tape price trends, drive replacement, electricity, request skew, caching, generator FLOPs, and future GPU price-performance improvements. For image generation, our analysis shows that prompt-based regeneration does not become cheaper than storage until around 2040, because every cache miss must still rerun the full prompt-to-artifact generation pipeline. We observe that widely used diffusion-based generation models operate in latent space, creating an alternative point in the cost tradeoff: instead of storing the final artifact or only the prompt, operators can store a compact intermediate representation (IR) and perform cheap on-demand decoding. Our analysis shows that caching combined with IR-based regeneration substantially reduces both storage and compute cost, making it at least 2x cheaper than both full-object storage and prompt-based regeneration even today. On a production image trace with 2.07 billion requests, the same conclusion holds: prompt-based regeneration is over 100x more expensive than storage, while IR-based regeneration reduces total cost to roughly half that of full-object storage while preserving interactive miss latency.

↑