arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

C-GAP:类感知与在线提示改进不均衡类上的视觉语言模型

C-GAP: Class-Aware and Online Prompting Improves Vision-Language Models on Imbalanced Classes

Francis Fernandez, Arash Jahangiri, Salimeh Sekeh

arXiv 2607.09008首次发表:更新:

发表机构

San Diego State University(圣地亚哥州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对安全关键感知系统检测小标签空间中罕见物体类别的问题,提出C-GAP框架,通过建立复合字幕基线并利用语言模型迭代细化提示,无需更新检测器权重,有效提升少数类检测精度,实验证明其成效显著。

AI 中文摘要

安全关键的感知系统必须在小标签空间中可靠地检测罕见物体类别,这是为数百个有密集注释的类别设计的长尾检测方法根本无法解决的。开放词汇检测器提供了一个有前景的替代方案,因为它们在推理时使用自然语言查询,使提示质量成为检测性能的首要杠杆。我们利用这一特性来解决类别不平衡问题:不是重新训练模型或收集额外注释,而是询问迭代细化输入到冻结检测器的语言提示是否能改善少数类检测。我们引入了C-GAP(字幕引导增强与提示),这是一个与检测器无关、无需注释的框架,分两个阶段运行。首先,我们建立一个将每个图像的场景描述与类数量上下文相结合的复合字幕基线,结果表明在多个开放词汇架构和基准测试中,该基线优于仅场景描述或仅类数量的提示。其次,一个语言模型迭代地单独细化每个图像的字幕,根据少数类AP@0.5与从复合基线得出的动态阈值,将试验分类为接受、暂定或重新生成桶。一旦获得足够的AP@0.5增益,细化就提前终止。在任何阶段都不更新检测器权重。我们的实验表明,C-GAP比基线将少数类平均精度提高了53%。在COCO上,C-GAP相对于复合基线将少数类AP@0.5提高了约81%(从17.69提高到32.09)。实验证实复合字幕为有效细化提供了关键基础:使用仅场景描述或仅类数量的提示作为细化起点收益递减,支持C-GAP的两个阶段作为必要贡献。

英文摘要

Safety-critical perception systems must reliably detect rare object classes within small label spaces, a setting that long-tailed detection methods, designed for hundreds of classes with dense annotation, fundamentally do not address. Open-vocabulary detectors offer a promising alternative, as they use natural language queries at inference time, making prompt quality a first-class lever for detection performance. We exploit this property to address class imbalance: rather than retraining models or collecting additional annotations, we ask whether iteratively refining the language prompts, fed to frozen detectors, can improve minority class detection. We introduce C-GAP Caption-Guided Augmentation and Prompting), a detector-agnostic, annotation-free framework that operates in two phases. First, we establish a composite caption baseline combining per-image scene descriptions with class-quantity context, which we show outperforms scene-description only or class-quantity-only prompts across multiple open-vocabulary architectures and benchmarks. Second, an LLM iteratively refines each image's caption individually, with trials triaged into accept, tentative, or regenerate buckets based on minority-class AP@0.5 against a dynamic threshold derived from the composite baseline. Refinement terminates early once sufficient AP@0.5 gain is achieved. No detector weights are updated at any stage. Our experiments shows that C-GAP improves minority-class average precision up to 53% over the baselines. On COCO, C-GAP improves minority-class AP@0.5 by ~81% relative over the composite baseline (17.69 -> 32.09). Experiments confirm that composite captions provide the critical foundation for effective refinement: using scene-description-only or class-quantity-only prompts as the refinement starting point yields diminishing returns, supporting both stages of C-GAP as necessary contributions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑