arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GVR-Coder:面向复杂文档与会议场景的结构化SVG生成视觉反馈框架

GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios

Yiming Xu, Jihua Kang, Chunsai Du, Qifan Zhang, Wangqiu Zhou, Yiting Wu, Tianqi Li, Qi Song

arXiv 2607.28073首次发表:更新:

发表机构

University of Science and Technology of China; ByteDance Inc.; Hefei University of Technology(中国科学技术大学; 字节跳动公司; 合肥工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对复杂文档与会议场景文本转SVG的三大挑战,提出GVR-Coder框架,引入定制数据集DocMeetSVG-100K,结合课程微调、双渲染反馈强化学习与智能体循环,生成质量优于基线。

AI 中文摘要

在严苛的专业环境与会议评审场景中,冗长文本常带来高认知负荷,为实现高效信息传递,将冗余文本转换为逻辑清晰的图表至关重要。可缩放矢量图形(SVG)因具备可编辑性与分辨率独立性,成为实现该目标的有效表示形式。然而,当前文本到SVG生成的研究仍面临三大核心挑战:(1)复杂、富含逻辑的图表数据集稀缺;(2)缺乏明确的布局先验,导致空间排列混乱;(3)缺少细粒度视觉反馈,无法验证渲染输出并修正美学缺陷。为应对这些挑战,在数据层面,我们引入DocMeetSVG-100K,这一专为文档创作与会议评审场景定制的大规模SVG数据集;在模型层面,我们提出GVR-Coder,一款用于从冗长专业文本生成高质量逻辑图表的新型框架。具体而言,我们采用课程驱动的拒绝采样微调,逐步提升模型对复杂结构的建模能力,同时在训练过程中明确融入布局约束知识。此外,我们引入双渲染反馈强化学习机制,该机制通过奖励信号提供隐式反馈,以联合优化结构复杂度与视觉美学。进一步地,我们设计生成-验证-修复智能体循环,通过明确的细粒度反馈与针对性优化提升生成质量。大量实验表明,GVR-Coder优于竞争性基线,可可靠生成逻辑连贯且视觉美观的图表。代码与数据可访问此https URL获取。

英文摘要

In demanding professional environments and meeting review scenarios, lengthy text often imposes a high cognitive load. To facilitate efficient information communication, transforming verbose text into logically clear diagrams is essential. Scalable Vector Graphics (SVG) provide an effective representation for this purpose due to their editability and resolution independence. However, current research on Text-to-SVG generation remains hindered by three major challenges: (1) the scarcity of datasets for complex, logic-rich diagrams; (2) the absence of explicit layout priors, which leads to chaotic spatial arrangements; and (3) the lack of fine-grained visual feedback to validate rendered outputs and correct aesthetic defects. To address these challenges, at the data level, we introduce DocMeetSVG-100K, a large-scale SVG dataset tailored for document authoring and meeting review scenarios. At the model level, we propose GVR-Coder, a novel framework designed to generate high-quality logical diagrams from lengthy professional texts. Specifically, we adopt a curriculum-driven rejection sampling fine-tuning to progressively enhance the model's capability in modeling complex structures, while explicitly incorporating layout constraint knowledge during training. In addition, we introduce reinforcement learning from dual rendering feedback, a mechanism that provides implicit feedback through reward signals to jointly optimize structural complexity and visual aesthetics. Furthermore, we design a generate-verify-repair agent loop, which improves generation quality through explicit, fine-grained feedback and targeted refinement. Extensive experiments demonstrate that GVR-Coder outperforms competitive baselines and reliably produces logically coherent and visually appealing diagrams. Code and data are available at https://github.com/CurryaNa/GVR-Coder.

DOI:10.1145/3767308.3835261

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑