AI 中文总结
该研究探究视觉生成中文本条件的缩放特性,采用GPG、ED量化结构化语言,据此构建结构化提示并训练提示生成器,所得系统在多基准上超越多数开放权重模型、匹配或超越最强封闭权重模型。
AI 中文摘要
我们研究视觉生成中文本条件的经验缩放特性。这类特性很少被测量,因为扩散损失不会随自然语言提示中的标记数量缩放。令人惊讶的是,我们发现收敛的扩散损失会随提示中结构化语言的量缩放。为量化结构化语言,我们采用两种互补度量:白盒似然度量(GPG)和黑盒属性度量(ED)。在受控训练运行中,收敛的扩散损失随GPG近似线性下降,且随ED遵循幂律。基于这些缩放特性,我们通过构建包含图像衍生语义和几何注释的结构化提示提升了\textit{diffusability},并通过监督微调、冷启动和验证器门控的在线策略蒸馏训练提示生成器提升了\textit{promptability}。所得系统在几乎所有组合、推理和世界知识基准上优于所有评估的开放权重模型,在多数评估中匹配或超越最强的封闭权重模型。
英文摘要
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled training runs, the converged diffusion loss decreases approximately linearly with GPG and follows a power law with ED. Guided by these scaling properties, we improve \emph{diffusability} by constructing structured prompts with semantic and geometric annotations derived from images, and improve \emph{promptability} by training a prompter through supervised fine-tuning, cold-start, and verifier-gated on-policy distillation. The resulting system outperforms all evaluated open-weight models on nearly every compositional, reasoning, and world-knowledge benchmark, while matching or surpassing the strongest closed-weight models on most evaluations.
CommentsCode: https://github.com/heheyas/context-scaling Models: https://huggingface.co/collections/heheyas/context-scaling Demo: https://heheyas-context-scaling.hf.space/ Project page: https://heheyas.github.io/context-scaling