发表机构
Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对A2UI声明式UI生成,探索了小模型成本下的前沿质量实现,研究了三类设计选择的影响,发现4B微调模型可在低成本下接近前沿模型质量,还给出了相关权衡与部署建议。
AI 中文摘要
诸如A2UI之类的声明式UI协议允许应用程序通过从目录中选择预构建组件并将其属性绑定到应用程序数据来生成交互式UI,而非从头编写前端代码。这种契约因安全性和一致性而对生产系统具有吸引力。一个悬而未决的问题是:低延迟、低成本的小模型能否达到基于A2UI的UI生成所需的质量?为回答此问题,我们系统研究了针对目录条件A2UI生成的三个可控设计选择:监督微调(SFT)数据构建方法、模型规模和组件目录规模。在两个React/TypeScript领域以及涵盖两个模型系列(Qwen 3.5 0.8B/2B/4B;SmolLM 3B)的四个基础检查点上,我们发现:(i)经过微调的4B学生模型在比前沿API调用成本低一个数量级以上的情况下,恢复了约98%的教师语义质量和约97%的教师视觉质量;(ii)两种增强策略(扰动目录和约束真实值)均在帕累托上优于未增强的全目录基线,同时在不同维度上表现出专业性;(iii)即使是小模型也能处理并受益于相对较大的组件目录规模。我们将这些结果提炼为面向从业者的权衡方案和这三个设计选择的部署建议。
英文摘要
Declarative UI protocols such as A2UI let applications generate interactive UIs by selecting pre-built components from a catalog and binding their props to application data, rather than emitting frontend code from scratch. This contract is attractive for production systems because of safety and consistency. An open question is: can low-latency and low-cost small models achieve the required quality for A2UI-based UI generation? To answer this, we systematically study three controllable design choices for catalog-conditioned A2UI generation: supervised fine-tuning (SFT) data construction method, model size, and component-catalog size. Across two React/TypeScript domains and four base checkpoints spanning two model families (Qwen 3.5 0.8B/2B/4B; SmolLM 3B), we find: (i) a 4B fine-tuned student recovers ~98% of teacher semantic quality and ~97% of teacher visual quality at more than an order of magnitude lower cost than frontier API calls; (ii) both augmented strategies (Perturbed-catalog and Constrained-GT) Pareto-dominate the unaugmented Full-catalog baseline, while specializing on different axes; (iii) even small models can handle and benefit from relatively large component catalog size. We distill these results into practitioner-facing trade-offs and deployment recommendations across the three design choices.