用于高分辨率目标检测的可变粒度分词(VGTok)
Variable-Granularity Tokenization for High-Resolution Object Detection
浏览论文内容
中文总结 AI 辅助
本文提出无需训练的可变粒度分词器VGTok,在航拍目标检测中固定分词预算,可在减少编码器计算和内存的同时,在VisDrone、AI-TOD-v2等数据集上创下最优检测性能。
中文摘要 AI 辅助
现有ViT目标检测器在所有学习阶段前会固定统一的分词网格,而原生分辨率的航拍目标检测器需在解析少像素目标与控制计算和内存限制间做取舍。本文提出VGTok,一种无需训练的分词器,可在编码器前从像素层面为每个区域设置块粒度:它通过多尺度形态学顶帽可分性(结合周围区域)对每个区域评分,再按每张图像的百分位数阈值处理这些评分,以此固定分词预算;结构张量门(λ_min)仅在二维目标结构支持的区域做细化,将一维杂波区域保持为粗粒度,最终得到的分词集是图像的严格划分。在搭载EVA-02 ViT-L编码器的Co-DETR检测器上,VGTok在40%至100%的所有分词预算下,均超过所有已发布的VisDrone验证集AP和AP_S指标:在40%预算时,它记录到44.22的AP,且在第一个Transformer块前丢弃了五分之三的序列;全分辨率时达到48.38的AP,比最强的已发布结果高出6.08。VGTok无需修改评分器和排名方式即可迁移至AI-TOD-v2数据集,在该数据集上达到37.27的AP和19.51的AP_vt,创下新的最优水平;作为直接插入冻结检查点的模块,它在78.5%的分词预算下达到36.29的AP,超过所有已发布结果,其中参数规模为3.763亿的检测器超越了参数量达30亿的多专家模型。研究表明,仅基于局部可分性和结构几何在骨干网络前固定的分词预算,能在航拍目标检测占主导的微小目标区域保持准确率,同时将编码器计算量降低3.1倍,内存占用降低1.9倍。代码和模型可在[http URL]和[http URL]获取。
英文摘要
ViT detectors fix a uniform token grid before any learned stage. A native-resolution aerial detector must then choose between resolving few-pixel objects and staying inside compute and memory limits. We introduce VGTok, a training-free tokenizer that sets patch granularity per region from pixels, ahead of the encoder. VGTok scores each region by multi-scale morphological top-hat separability from its surround, then thresholds those scores at a per-image percentile, which fixes the token budget. A structure-tensor gate ($λ_{\min}$) refines only where two-dimensional object structure supports it, leaving one-dimensional clutter coarse. The resulting token set is a strict partition of the image. In a Co-DETR detector with an EVA-02 ViT-L encoder, VGTok clears every published VisDrone-val AP and AP$_S$ at every budget from 40\% to 100\% of tokens. At 40\% it records 44.22 AP with three fifths of the sequence discarded before the first transformer block; dense, it reaches 48.38 AP, $6.08$ above the strongest published entry. VGTok transfers to AI-TOD-v2 untouched, same scorer and same rank, and sets a new state of the art at 37.27 AP and 19.51 AP$_{vt}$. As a pure drop-in into a frozen checkpoint it reaches 36.29 AP at 78.5\% of tokens, above every published entry, where our 376.3M-parameter detector clears a 3.0B multi-expert model. We show that a token budget fixed before the backbone, from local separability and structure geometry alone, holds accuracy on the tiny-object regimes that dominate aerial detection, at $3.1\times$ less encoder compute and $1.9\times$ less encoder memory. Code and models are available at \href{https://github.com/khayrulbuet13/vgtok}{\texttt{github.com/khayrulbuet13/vgtok}} and \href{https://huggingface.co/khayrulbuet13/vgtok}{\texttt{huggingface.co/khayrulbuet13/vgtok}}.
发表机构
- Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室)
机构由 AI 辅助整理,请以论文原文为准。