arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21460cs.CVcs.AI

FigmaTrace:捕捉人类Figma设计工作流程中的创意细节

FigmaTrace: Capturing Creative Nuances in Human Figma Design Workflows

Darshan Deshpande, Yoshinari Fujinuma, Martyna Markiewicz, Devanshu Bansal, Shivani Jain, Nicholas Saban, Chirag Maheshwari, Anand Kannappan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究构建了FigmaTrace数据集,训练模型后其在四个分布外智能体GUI环境中性能可与前沿闭源模型媲美,且通过定性分析关联了性能提升与数据集的趋势,同时开源了数据集和最佳模型。

中文摘要 AI 辅助

视觉语言模型近期在目标检测等多个客观可验证领域取得了进展,但在主观且具有创造性的设计任务上表现仍不佳。这一性能差距的主要原因是缺乏高质量的人类工作流程数据,这类数据需涵盖各类偏好与决策,而这些正是人类专家擅长设计任务的关键。本研究中,我们首先定义了一套由专家精心整理的独特设计技能与最佳实践分类法,并将其扩展为126项开放式、主观且长期的任务。基于该分类法及专家解决方案,我们的数据集FigmaTrace包含超过200小时的人类捕获视频数据,通过一种新颖的基于设计阶段的方法将其转换为3469条设计轨迹。我们使用该数据集训练了四个模型,结果显示,在四个分布外的智能体图形用户界面(GUI)环境中,基于FigmaTrace训练的模型取得了可与前沿闭源模型(如\textsc{Claude-Opus-5}和\textsc{GPT-5.6-Sol})相媲美的性能提升。我们进一步开展了 ablation( ablation即消融实验,是机器学习中用于评估模型组件重要性的实验方法),将这些性能提升归因于基于设计阶段的视频到轨迹转换,该方法优于先前基于长度的转换方法。最后,我们对性能最佳的\textsc{Qwen3.8-27B}模型输出进行了定性分析,以更好地将性能提升与FigmaTrace的趋势相关联。我们将数据集和最佳模型开源供社区使用。

英文摘要

Vision Language Models have recently shown improvements in several objective and verifiable domains such as object detection but continue to underperform on subjective and creative design tasks. A major contributor to this performance gap is the lack of high quality human workflow data that captures a diverse set of preferences and decisions that make human experts good at design tasks. In this work, we first define a unique, expert curated taxonomy of design skills and best practices which we further expand into a set of 126 open ended, subjective, long horizon tasks. Built on top of this and expert solutions, our dataset FigmaTrace contains over 200 hours of human captured video data converted into 3469 design trajectories using a novel design phase-based method. We use our dataset to train four models and show that training on FigmaTrace leads to a performance improvement comparable to frontier closed models such as \textsc{Claude-Opus-5} and \textsc{GPT-5.6-Sol} on four out of distribution agentic GUI environments. We further perform a useful ablation to attribute these performance improvements to a design phase-based video to trajectory conversion which outperforms prior length-based conversion approaches. Finally, we perform a qualitative analysis on the best performing \textsc{Qwen3.8-27B} outputs to better correlate performance improvements to FigmaTrace's trends. We open source our dataset and the best model for the community.

发表机构

  • Patronus AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑