发表机构
Zhejiang University; Taobao & Tmall Group of Alibaba(浙江大学; 阿里巴巴淘宝天猫集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对工业UI代码生成中MLLMs超越视觉保真度的需求,提出TaoD2C-Bench基准,包含2,861个设计及97,652条注释,定义三个任务,评估八个模型,发现显著差距,并指出视觉重建不等于满足实现要求。
AI 中文摘要
多模态大语言模型(MLLMs)的一个关键挑战是超越视觉识别,实现约束感知的跨模态推理。这涉及将视觉线索与其他模态的信息相结合,以理解在特定领域规则下元素之间的关系。这一挑战在工业设计到代码(D2C)转换中尤为明显,该过程将用户界面(UI)设计转换为代码,要求MLLMs将设计图像与杂乱无章的图层元数据连接起来,推断组件和布局的实现要求,并在目标库约束下将其实现为代码。然而,这些能力在现实工业环境中尚未得到充分评估。为填补这一空白,我们提出了TaoD2C-Bench,一个用于评估MLLMs在工业应用中生成满足实现要求的UI代码能力的基准。TaoD2C数据集包含来自17个商业平台的2,861个生产设计,以及涵盖四个类别(组件、分组、对齐和位置)的97,652条专家注释。这些注释区分了必需约束与允许的实现选择。TaoD2C-Bench定义了三个任务:端到端UI代码生成、需求推断和需求实现。对八个MLLMs的评估揭示了在生成满足实现要求的UI代码方面存在显著差距,同时推断和实现方面表现出不同的性能特征。我们进一步表明,MLLMs的视觉重建能力并不必然意味着生成满足这些要求的代码的能力。我们发布TaoD2C以支持工业UI代码生成的研究。
英文摘要
A key challenge for multimodal large language models (MLLMs) is moving beyond visual recognition to constraint-aware cross-modal reasoning. This involves combining visual cues with information from other modalities to understand elements' relationships under domain-specific rules. This challenge is acutely evident in industrial design-to-code (D2C), which converts user interface (UI) designs into code and requires MLLMs to connect design images with disorganized layer metadata, infer component and layout implementation requirements, and realize them in code under target-library constraints. However, these capabilities remain insufficiently evaluated in realistic industrial settings. To fill this gap, we present TaoD2C-Bench, a benchmark for evaluating MLLMs' ability to generate UI code that satisfies implementation requirements in industrial applications. The TaoD2C dataset consists of 2,861 production designs from 17 commercial platforms with 97,652 expert annotations across four categories: Component, Group, Alignment, and Position. These annotations distinguish required constraints from permitted implementation choices. TaoD2C-Bench defines three tasks: end-to-end UI code generation, requirement inference, and requirement realization. Evaluating eight MLLMs reveals substantial gaps in generating UI code that satisfies implementation requirements, alongside distinct performance profiles in inference and realization. We further show that MLLMs' visual reconstruction ability does not necessarily imply an ability to generate code that meets these requirements. We release TaoD2C to support research on industrial UI code generation.
Comments26 pages