发表机构
Dongguk University(东国大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对YOLO-World难以处理模糊提示的问题,提出Vague2Detect混合流水线,结合微调Sentence-BERT和知识库检索候选,GPT-3.5扩展知识库,将VPSR从32%提升至61%,最高达85%。
AI 中文摘要
现实世界中的检测器通常必须解释功能性或模糊的提示,然而诸如YOLO之类的传统模型仍局限于固定的类别列表。即使是像YOLO-World这样的开放词汇模型,也经常无法将模糊的语言与预期的对象正确对齐。基于我们先前的工作“使用LLMs和视觉-语义匹配的常识引导开放世界目标检测”,我们解决了YOLO-World在基于任务查询定位方面的局限性。我们提出了Vague2Detect,一种混合流水线,其中微调的Sentence-BERT从结构化的家庭知识库(KB)中检索候选对象,而YOLO-World验证它们在图像中的存在。对于知识库之外的提示,大型语言模型(GPT-3.5-turbo)生成候选描述,动态扩展知识库以涵盖新概念。在使用自定义图像和Open Images V7子集的家居场景基准测试中,仅使用YOLO-World的模糊提示成功率(VPSR)仅为32%,即能够将模糊查询映射到正确检测的能力。相比之下,Vague2Detect将性能提升至61%的VPSR且具有高精度,并在使用GPT回退时最高可达85%。
英文摘要
Real-world detectors must often interpret functional or ambiguous prompts, yet conventional models such as YOLO remain restricted to fixed class lists. Even open-vocabulary models like YOLO-World frequently misalign vague language with the intended objects. Building on our prior work Commonsense-Guided Open-World Object Detection Using LLMs and Visual-Semantic Matching, we address YOLO-World's limitations in grounding task-driven queries. We propose Vague2Detect, a hybrid pipeline in which a fine-tuned Sentence-BERT retrieves candidates from a structured household Knowledge Base (KB), and YOLO-World verifies their presence in the image. For prompts outside the KB, a large language model (GPT-3.5-turbo) generates candidate descriptions, dynamically expanding the KB to cover novel concepts. On a benchmark of household scenes using custom images and an Open Images V7 subset, YOLO-World alone achieves only 32% Vague Prompt Success Rate (VPSR), the ability to map ambiguous queries to correct detections. In contrast, Vague2Detect improves performance to 61% VPSR with high precision, and up to 85% when augmented with GPT fallback.
Comments15 pages, 4 figures, 3 tables. Code: https://github.com/ibrohimgets/Vague2Detect