arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.07577cs.CVcs.AIcs.RO

开放世界层次感知:针对类无关候选区域的分类抽象,用于安全处理词汇外道路对象

Open-World Hierarchical Perception: Taxonomic Abstraction over Class-Agnostic Proposals for the Safe Handling of Out-of-Vocabulary Road Objects

Felix Schaller

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出开放世界层次感知层,在类无关区域候选上实现分类抽象,通过结合三类开放世界信号,在自动驾驶留类基准测试中,可安全处理词汇外道路对象且无分类错误。

中文摘要 AI 辅助

自动驾驶的闭集检测器必须为每个对象分配固定标签集中的某一标签。对于该集合外的对象(如马车、道路碎片、乡村道路上的牲畜),它要么只能强行赋予一个置信度高但错误的具体标签,要么直接丢弃该对象。该系列的前期工作用层次分类法和运行时抽象规则取代了平面标签集,但仅在闭集检测器已生成的边界框上进行了评估。本文将层次抽象扩展到类无关区域候选之上,使得闭集检测器从未框选的对象仍能被分类或标记;本文报告了一项对三种开放世界信号(类无关分割、基于外观的分布外评分、单目深度)的可行性研究,表明为何没有单一二维线索足够,以及它们如何组合;本文还开展了前期论文无法进行的评估:针对真实标注对象的留类基准测试。在留出7个COCO类别并对其235个真实裁剪区域进行分类时,平面闭集头部100%会发出一个置信度高的错误具体标签(其中37%属于错误的超类别,例如将动物命名为车辆),而层次层则零置信度错误具体标签,且安全处理了94%的对象(正确超类别或明确的UNKNOWN OBSTACLE)。本文明确指出这是一项安全结果而非特异性结果:仅26%的时间能恢复正确超类别,剩余69%被保守标记为未知。本文的贡献是一种开放世界感知层,该层从未对词汇外对象产生置信度高的分类错误,同时如实说明了其代价。

英文摘要

A closed-set detector for autonomous driving must assign every object one of a fixed set of labels. On an object outside that set (a horse-drawn carriage, road debris, livestock on a rural road) it can only force a confident but wrong specific label or drop the object. Prior work in this series replaced the flat label set with a hierarchical taxonomy and a runtime abstraction rule, but evaluated it only on the boxes a closed detector already produces. This paper takes the layer open-world: we place taxonomic abstraction on top of class-agnostic region proposals so objects the closed detector never boxes can still be classified or flagged; we report a feasibility study of three open-world signals (class-agnostic segmentation, appearance-based out-of-distribution scoring, monocular depth) that shows why no single 2D cue suffices and how they compose; and we run the evaluation the earlier papers could not, a ground-truth leave-classes-out benchmark on real annotated objects. Holding out seven COCO classes and classifying their 235 ground-truth crops, a flat closed head emits a confident wrong specific label 100% of the time (37% of them in the wrong super-category, e.g. an animal named as a vehicle), whereas the hierarchical layer emits zero confident wrong specific labels and safely handles 94% of the objects (a correct super-category, or an explicit UNKNOWN OBSTACLE). We are explicit that this is a safety result, not a specificity one: the correct super-category is recovered only 26% of the time and the remaining 69% are conservatively flagged unknown. The contribution is an open-world perception layer that never makes a confident categorical mistake on an out-of-vocabulary object, together with an honest account of its cost.

补充信息

↑