RegionDet:超越物体实例的区域检测基准
RegionDet: A Benchmark for Region Detection Beyond Object Instances
浏览论文内容
中文总结 AI 辅助
针对现有基准无法评估非物体实例的区域目标的问题,本文提出区域检测任务,构建RegionDet基准并评估不同检测器,揭示视觉-语言检测器的以物体为中心偏差,分析区域检测的关键挑战。
中文摘要 AI 辅助
物体检测是计算机视觉中的基础任务,通过定位离散且边界清晰的物体实例,在标准基准上已取得显著进展。但现实场景中许多视觉目标并非单个物体,而是由视觉状态、场景上下文、物体关系及人类活动定义的区域,例如施工区域、受损道路区域、队列、群体交谈、商贩区域等。现有检测基准主要围绕物体实例构建,对这类区域目标的系统评估支持有限。为解决这一缺口,本文提出区域检测任务,将传统物体检测扩展至物体实例之外,并构建区域目标定位基准RegionDet。RegionDet包含8个区域类别,包括施工、路口、损伤、排队、交谈、商贩、等待、行走,采用COCO风格的边界框标注与评估协议。本文在RegionDet上系统评估了代表性的闭集检测器及零样本/开放词汇检测器,结果显示闭集检测器可在监督下部分学习区域级模式,而零样本/开放词汇检测器表现极差,揭示了当前视觉-语言检测器强烈的以物体为中心的偏差。进一步分析凸显了区域检测的关键挑战,包括边界线索薄弱、上下文依赖性强、关系级区域理解不足。RegionDet将被发布。
英文摘要
Object detection is a fundamental task in computer vision and has achieved remarkable progress on standard benchmarks by localizing discrete and well-bounded object instances. However, many visual targets in real-world scenarios are not individual objects, but regions defined by visual states, scene context, object relations, and human activities, such as construction areas, damaged road regions, queues, group conversations, and vendor regions. Existing detection benchmarks are mainly built around object instances, providing limited support for systematically evaluating such region targets. To address this gap, we introduce Region Detection, a task that extends conventional object detection beyond object instances, and construct RegionDet, a benchmark for region target localization. RegionDet contains eight region categories, including Construction, Crossing, Damage, Queuing, Talking, Vendor, Waiting, and Walking, with COCO-style bounding-box annotations and evaluation protocols. We systematically evaluate representative closed-set and zero-shot/open-vocabulary detectors on RegionDet. Results show that closed-set detectors can partially learn region-level patterns under supervision, while zero-shot/open-vocabulary detectors struggle severely, revealing the strong object-centric bias of current vision-language detectors. Further analyses highlight key challenges in Region Detection, including weak boundary cues, strong context dependency, and insufficient relation-level region understanding. The RegionDet will be released.