arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06116cs.HCcs.CV

MM-SVGEdit:面向UI设计的多模态驱动SVG编辑

MM-SVGEdit: A Multimodal-Driven SVG Editing for UI Design

  • Beihang University(北京航空航天大学)

机构由 AI 辅助整理,请以论文原文为准。

Shibo Yang, Yuqing Gao, Zipeng Liu

AI总结:

针对传统SVG编辑门槛高和LLM方法可控性差的问题,提出多模态驱动的MM-SVGEdit,采用视觉定位加修改的两阶段策略,支持指令和直接操作,在自建数据集上显著提升编辑准确率、效率和用户控制感,并降低资源消耗。

AI中文摘要:

在UI设计领域,可缩放矢量图形(SVG)被广泛用作设计媒介。然而,传统的SVG编辑技术入门门槛高,且需要繁琐的手动迭代,而基于大语言模型(LLM)的编辑解决方案则存在准确性低和用户可控性差的问题。为解决这些问题,我们提出了MM-SVGEdit,一种多模态驱动的SVG编辑方法,它整合了传统SVG编辑与基于LLM的方法。我们引入了一个两阶段策略,即先进行视觉定位,再进行修改。两个阶段都支持两种交互模态:自然语言指令和直接操作(鼠标和键盘)。我们在一个自构建的、从UI生成的14,476个问答对数据集上训练并评估了MM-SVGEdit,该数据集涵盖了针对单个和多个UI目标的11种编辑操作类型。结果表明,MM-SVGEdit提高了SVG编辑的准确性、效率和用户感知的控制力,同时减少了令牌消耗和响应时间。

英文摘要:

In the field of UI design, Scalable Vector Graphics (SVG) is widely used as a design medium. However, traditional SVG editing techniques have high entry barriers and require cumbersome manual iteration, while LLM-based editing solutions suffer from low accuracy and poor user controllability. To address these issues, we propose MM-SVGEdit, a multimodal-driven SVG editing approach that integrates traditional SVG editing and LLM-based methods. We introduce a two-stage strategy in which visual grounding is followed by modification. Both stages support two interaction modalities: natural language instructions and direct manipulation (mouse and keyboard). We trained and evaluated MM-SVGEdit on a self-constructed dataset of 14,476 question-answer pairs generated from UIs, covering 11 types of editing operations on both single and multiple UI targets. The results show that MM-SVGEdit improves SVG editing accuracy, efficiency, and user-perceived control while reducing token consumption and response time.

↑