AI 中文总结
研究旨在解决桌面界面窗口组件检测问题,提出基于计算机视觉的TargetFinder系统,利用微调YOLO网络和新数据集,实现轻量级屏幕监控与低延迟检测,性能优于基线方法,还展示了通用目标感知技术可行性并开源相关资源。
AI 中文摘要
“目标感知”指向技术,如气泡光标或语义指向,通过利用目标位置知识优于传统指向。然而,缺乏与应用无关的窗口组件几何信息限制了它们在桌面上的应用。我们提出了TargetFinder,一个基于计算机视觉的用于实时检测GUI窗口组件的系统。它利用在包含520个带注释桌面截图(约38000个注释)的新数据集上训练的多个微调YOLO网络,涵盖Windows、macOS、Ubuntu和网页界面。TargetFinder使用轻量级屏幕监控和低延迟检测,实现了适合交互使用的毫秒级响应。评估表明TargetFinder优于基线方法,而气泡光标和语义指向的系统范围实现证明了部署跨应用通用目标感知技术的可行性。我们发布了数据集、模型、注释工具和开源库用于研究和应用。
英文摘要
''Target-aware'' pointing techniques, like Bubble Cursor or Semantic Pointing, outperform traditional pointing by leveraging knowledge of target locations. Yet the lack of application-agnostic widget geometry information limits their adoption across the desktop. We present TargetFinder, a computer vision-based system for real-time detection of GUI widgets. TargetFinder leverages several fine-tuned YOLO networks trained on a new dataset of 520 annotated desktop screenshots (~38,000 annotations) spanning Windows, macOS, Ubuntu, and web interfaces. TargetFinder uses lightweight screen monitoring and low-latency detection, achieving millisecond responsiveness suitable for interactive use. Evaluations show that TargetFinder outperforms the baseline methods (OmniParser and REMAUI), while system-wide implementations of Bubble Cursor and Semantic Pointing demonstrate the feasibility of deploying universal target-aware techniques that work across applications. We release the dataset, models, annotation tool, and an open-source library for research and applications.