arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22916cs.CV

规划与渲染协同:自回归布局与扩散的深度融合用于视觉文本生成

Planning and Rendering in Concert: DeepFusion of Autoregressive Layouts and Diffusion for Visual Text Generation

Guanqiao Chen, Jingru Tan, Dongxing Mao, Catherine Chen, Zijian Du, Libo Qin, Hu Jian Guo, Alex Jinpeng Wang

首次发表
浏览论文内容

中文总结 AI 辅助

DuetGen通过深度融合自回归规划与扩散渲染,联合学习文本布局与图像生成,在CVTG-2K和LongText-Bench上以较小模型接近Qwen-Image性能,实现高效自主视觉文本生成。

中文摘要 AI 辅助

从提示生成富含文本的图像需要文本保真度以及将文本连贯地整合到周围图像中。显式布局可以提供关于应出现什么文本及其位置的结构化指导,但仅凭良好的规划并不能保证渲染器会忠实地实现它。现有的基于布局的自回归-扩散系统通常分别优化规划和渲染,阻止规划器的表示与图像合成联合调整。我们引入了DuetGen,一个基于DeepFusion构建的自主视觉文本生成器,它联合学习自回归规划和连续扩散渲染。DeepFusion将扩散变换器条件化于规划器的提示和边界框-内容隐藏状态,允许渲染监督塑造连接文本规划与视觉输出的表示。其联合目标结合了自回归规划监督、文本区域加权扩散学习和辅助坐标监督,以维持结构化规划、强调文本承载区域并提高规划器表示的空间精度。在推理期间,相位感知注意力调制加强了图像区域与其匹配的坐标和内容状态之间的对应关系,促进生成规划的区域特定执行。凭借2B规划器和4B单流DiT,DuetGen在CVTG-2K上达到0.8293的词准确率,在LongText-Bench上达到0.938的准确率,在两个基准上紧密匹配规模大得多的Qwen-Image。这些结果证明了联合学习的规划表示和区域特定渲染对自主视觉文本生成的价值。

英文摘要

Generating text-rich images from prompts requires both textual fidelity and the coherent integration of text into the surrounding image. An explicit layout can provide structured guidance about what text should appear and where, but a well-formed plan alone does not guarantee that the renderer will realize it faithfully. Existing layout-based AR-diffusion systems typically optimize planning and rendering separately, preventing the planner's representations from being adapted jointly with image synthesis. We introduce DuetGen, an autonomous visual text generator built on DeepFusion, which jointly learns autoregressive planning and continuous diffusion rendering. DeepFusion conditions a diffusion transformer on the planner's prompt and bbox-content hidden states, allowing rendering supervision to shape the representations connecting textual plans with visual outputs. Its joint objective combines autoregressive plan supervision, text-region-weighted diffusion learning, and auxiliary coordinate supervision to maintain structured planning, emphasize text-bearing regions, and improve the spatial precision of planner representations. During inference, Phase-Aware Attention Modulation strengthens the correspondence between image regions and their matched coordinate and content states, facilitating region-specific execution of the generated plan. With a 2B planner and a 4B single-stream DiT, DuetGen achieves 0.8293 word accuracy on CVTG-2K and 0.938 accuracy on LongText-Bench, closely matching the substantially larger Qwen-Image on both benchmarks. These results demonstrate the value of jointly learned planning representations and region-specific rendering for autonomous visual text generation.

发表机构

  • University of Science and Technology of China(中国科学技术大学)
  • Central South University(中南大学)
  • Louisiana State University(路易斯安那州立大学)
  • Arizona State University(亚利桑那州立大学)
  • Sun Yat-sen University(中山大学)

机构由 AI 辅助整理,请以论文原文为准。

↑