发表机构
Federal University of Paraná; Paraná Military Police(巴拉那联邦大学; 巴拉那宪兵队)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对基础模型在未知数据分类中缺乏细粒度表征的问题,提出VeriCam管道,结合验证任务训练的图像模型与Leiden聚类算法,在LPLCv2跨设备场景中取得93.45 F1值、80.13 V-Measure值的良好性能,构建了公平无偏的未知数据分类基准。
AI 中文摘要
基础模型的出现开启了零样本分类的新时代,但仍存在关键挑战。尽管基础模型凭借海量预训练知识具备出色的泛化能力,然而图像、文本类基础模型以及视觉-文本混合模型,均缺乏现实世界部分任务所需的基于细微差别的细粒度类别分离表征能力。为解决现有研究的不足,本文提出VeriCam,这一管道旨在学习高度专业化的特征,以实现对未见数据中未知类别的分类。VeriCam利用为验证任务训练的图像模型的表征能力,这类模型会形成包含细粒度细节的复杂特征空间;通过训练模型区分同一类别与不同类别的图像对,构建表征数据点间类别关系的关系图。本文提出两种图聚类方法:一种是朴素算法,另一种是针对Leiden图聚类算法的特定设置。该管道在LPLCv2数据集上进行验证,该数据集包含真实世界的交通监控图像。研究发现该数据集存在固有的采集设备偏差,这对下游车牌识别任务(如OCR)构成泛化挑战;为此,本文采用与标签无关的方法动态识别采集设备,从而构建公平无偏的基准。在跨设备场景中,该管道在验证基准中达到93.45的F1分数,在聚类步骤中达到80.13的V-Measure分数,所有代码均公开提供。
英文摘要
The advent of foundation models have enabled a new era in zero-shot classification. Yet, key challenges persist. Despite their impressive generalization power that leverages the immense pre-training knowledge, both foundation models for image and text as well as vision-text hybrids lack the representational power needed for fine-grained, minutiae-based class separation that some real-world tasks require. To address the current gaps in the literature, we propose VeriCam, a pipeline designed to learn highly specialized features that enable classification of unknown classes in unseen data. VeriCam works by leveraging the representation power of image models trained for the verification task, where the model develops an intricate feature space that incorporates fine-grained details. By training a model to discriminate between pairs of images from the same and different classes, a relational graph is constructed, representing the class relationships between data points. We then present two approaches for graph clustering: a naive algorithm and a specific setup for the Leiden graph clustering algorithm. The pipeline is validated on the LPLCv2 dataset, which comprises real-world traffic surveillance images. We show that the dataset carries an inherent capture device bias that is posed as a generalization challenge for downstream License Plate recognition tasks such as OCR. As such, we dynamically identify capture devices with a label-agnostic approach, enabling the construction of a fair and unbiased benchmark. In the cross-device scenario, our pipeline reaches an F1-Score of 93.45 in the verification baseline and a V-Measure score of 80.13 in the clustering step. All code is publicly available at https://github.com/lmlwojcik/VeriCam
CommentsSIBGRAPI WIP 2026