作为自演进UI转代码生成视觉修复上下文的评分标准
Rubrics as Visual-Repair Context for Self-Evolving UI-to-Code Generation
浏览论文内容
中文总结 AI 辅助
该研究针对视觉语言模型UI转代码自演进不稳定的问题,提出RubSE框架,利用评分标准作为视觉修复上下文,在多模型多基准上显著提升了生成性能与稳定性。
中文摘要 AI 辅助
大型视觉语言模型在UI转代码生成领域已取得显著进展,但其测试阶段的自演进过程仍不稳定。我们首先识别出一个名为视觉修复耦合的根本障碍:局部代码编辑可能会通过布局、样式和组件依赖关系传播,在修正一处视觉不匹配的同时,破坏原本符合要求的区域。为解决该问题,我们提出了RubSE,即评分标准引导的自演进框架,该框架使用评分标准将视觉反馈表示为结构化的视觉修复上下文。在每一轮迭代中,RubSE会生成带类型的候选评分标准,选择一个优先级最高的修复目标,并将之前选中的评分标准存储为历史记录,从而引导每一次修订朝着范围明确的视觉修复方向进行,同时抑制重复或范围过宽的修改。在六个视觉语言模型和三个UI转代码基准上的评估表明,RubSE在最终轮和最佳轮设置下均显著优于朴素自演进方法,实现了更稳定的迭代轨迹和更高的轨迹级性能上限。进一步分析显示,RubSE通过提升从严重视觉退化中恢复的能力来缓解轨迹崩溃问题,且更强的评分标准生成器能够将有效的视觉修复指导传递给较弱的代码改进器。
英文摘要
Large vision-language models have shown strong progress in UI-to-code generation, yet their test-time self-evolution remains unstable. We first identify a fundamental obstacle, termed visual repair coupling: a local code edit may propagate through layout, style, and component dependencies, correcting one visual mismatch while degrading regions that were previously faithful. To address this issue, we present RubSE, a Rubric-guided Self-Evolution framework that uses rubrics to represent visual feedback as a structured visual-repair context. At each refinement round, RubSE generates typed candidate rubrics, selects one prioritized repair target, and stores previously selected rubrics as history, thereby steering each revision toward a well-scoped visual repair while discouraging repeated or over-broad changes. Evaluations across six VLMs and three UI-to-code benchmarks demonstrate that RubSE substantially outperforms naïve self-evolution in final-round and best-round settings, achieving more stable refinement trajectories and a higher trajectory-level performance ceiling. Further analysis shows that RubSE mitigates trajectory collapse by improving recovery from severe visual regressions, and that stronger rubric generators can transfer effective visual-repair guidance to weaker code improvers.
发表机构
- University of Maryland(马里兰大学)
- Microsoft(微软公司)
机构由 AI 辅助整理,请以论文原文为准。