Moonworks Lunara:建模艺术智能
Moonworks Lunara: Modeling Artistic Intelligence
浏览论文内容
中文总结 AI 辅助
Moonworks Lunara 提出艺术智能为探索驱动的世界实现,采用扩散混合 Transformer 架构,在美学质量上超越多个模型,以亚10B参数和亚10秒延迟推进视觉智能。
中文摘要 AI 辅助
我们将“艺术智能”定义为由探索驱动的世界实现,在保留必须真实的语义、艺术和构图结构的同时,为创造性可能性留出空间。Moonworks Lunara 是一个文本到图像模型,通过一种新颖的扩散混合 Transformer 架构实现了这一框架。一种新的训练算法通过信息丰富的样本采集和针对性地注入人类创作的艺术,迭代地演化数据分布。我们将 Lunara 与七个图像生成模型进行基准测试,包括 FLUX.2-Klein-4B、Qwen-Image(20B)和 GPT-Image-1-Mini。以 GPT-5.6 Sol 作为评估器,Lunara 在“美学质量”方面排名第一(8.473 对比 GPT-Image-1-mini 的 8.457),在“情感共鸣”方面排名第二,并在“内容完整性”方面保持竞争力。对相同评估集进行的盲人人工评估证实了自动指标,将 Lunara 排名第一。在包括 CLIPScore 和 LAION Aesthetic Predictor 在内的传统指标下,它也是最强模型之一。在 GenEval 上,Lunara 与更广泛的 16 个模型(包括 GPT Image 2 和 Seedream 4.0)相比取得了有竞争力的性能。这些结果使 Lunara 在保持低于 10B 活跃参数规模和低于 10 秒推理延迟的同时,处于艺术智能的前沿。Lunara 通过将问题从模型能否正确生成图像转变为它们能多深入地解读意义并将其实现为富有想象力和表现力的世界,推进了通用视觉智能的前沿。
英文摘要
We formulate \emph{Artistic Intelligence} as exploration driven world realization, leaving space for creative possibility while preserving the semantic, artistic, and compositional structure that must remain true. Moonworks Lunara, a text-to-image model, implements this framework with a novel Diffusion Mixture Transformer architecture. A new training algorithm iteratively evolves the data distribution through informative sample acquisition and targeted injection of human-created art. We benchmark Lunara against seven image-generation models, including FLUX.2-Klein-4B, Qwen-Image (20B), and GPT-Image-1-Mini. With GPT-5.6 Sol as evaluator, Lunara ranks first in \emph{Aesthetic Quality (8.473 vs. 8.457 GPT-Image-1-mini)}, second in \emph{Emotional Resonance}, and remains competitive in \emph{Content Integrity}. A blind human evaluation over the same evaluation set corroborates the automated metrics, ranking Lunara first. It also stays among the strongest models under conventional measures including CLIPScore and LAION Aesthetic Predictor. On GenEval, Lunara achieves competitive performance against a broader set of 16 models, including GPT Image 2 and Seedream 4.0. These results place Lunara at the frontier with Artistic Intelligence while maintaining a sub-10B active-parameter footprint and sub-10-second inference latency. Lunara advances the general visual intelligence frontier by shifting the question from whether models can get images right to how deeply they can interpret meaning and realize it as imaginative, expressive worlds.
发表机构
- Moonworks
机构由 AI 辅助整理,请以论文原文为准。