Trifuse: Enhancing Attention-Based GUI Grounding via Multimodal Fusion
Trifuse: 通过多模态融合增强基于注意力的GUI定位
机构 * College of Computer Science and Technology, National University of Defense Technology(计算机科学与技术学院,国防科技大学) ; Intelligent Game and Decision Lab, Academy of Military Sciences(智能游戏与决策实验室,军事科学院)
专题命中 多模态训练与对齐 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI
AI总结 Trifuse通过多模态融合提升GUI定位性能,整合注意力、OCR文本和图标描述语义,无需任务微调即可实现高效定位。
Comments 17 pages, 10 figures