融合UI结构与语义的面向特征的应用屏幕检索与聚类
Fusing UI Structure & Semantics for Feature-Oriented App Screen Retrieval & Clustering
浏览论文内容
中文总结 AI 辅助
该研究针对UI编程中代码与图形表示的抽象鸿沟,提出多模态神经符号嵌入技术FRAME,在三项基准测试中提升了屏幕检索与聚类的性能,可增强自动化UI设计与测试工具。
中文摘要 AI 辅助
用户界面(UI)编程因代码与图形化软件表示之间存在复杂的抽象鸿沟而颇具挑战性。为弥合这一鸿沟,UI编程工具常依赖屏幕检索与聚类,而这需要基于重叠特征的准确相似度度量。然而,计算面向特征的相似度颇具难度,因为功能相似的屏幕往往存在设计差异。为解决该问题,我们提出FRAME(基于图形结构理解的强化用户界面屏幕嵌入,ReinForced UseR InterfAce Screen EMbedding with Graphical Structural ComprEhension),这是一种多模态神经符号嵌入技术。FRAME构建UI组件的符号化图表示,以编码显著关系并捕捉不同屏幕间的特征模式。它利用大型视觉语言模型进行视觉与词汇编码,同时采用一种新型UI特定计算几何算法实现加权嵌入传播。在三个基准测试中,FRAME在搜索任务上较强基线模型的MRR(平均 reciprocal rank)提升最高达13%,在聚类准确率上提升7.6个百分点。全面的消融研究进一步证实了各组件的效用,证明FRAME具备增强自动化UI设计与测试工具的潜力。
英文摘要
User Interface (UI) programming is challenging due to the complex abstraction gap between code and graphical software representations. To bridge this gap, UI programming tools often rely on screen retrieval and clustering, which require accurate similarity measures based on overlapping features. However, computing feature-oriented similarity is difficult because screens with similar functionality often exhibit design variations. To address this, we propose FRAME (ReinForced UseR InterfAce Screen EMbedding with Graphical Structural ComprEhension), a multi-modal, neuro-symbolic embedding technique. FRAME constructs symbolic, graph-based representations of UI components to encode salient relationships and capture feature patterns across different screens. It leverages large vision-language models for visual and lexical encoding, alongside a novel UI-specific computational geometry algorithm that enables weighted embedding propagation. Across three benchmarks, FRAME outperforms strong baselines by up to 13% MRR in search and 7.6 percentage points in clustering accuracy. A comprehensive ablation study further confirms the benefit of each component, demonstrating FRAME's potential for enhancing automated UI design and testing tools.