arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11770cs.CVcs.DCcs.LG

在实时检测流水线中实现近零开销的多模型分层分类

Achieving Near-Zero-Overhead Multi-Model Hierarchical Classification in Real-Time Detection Pipelines

Vaishnav Raju

首次发表
浏览论文内容

中文总结 AI 辅助

针对边缘视觉系统检测分类流水线的串行瓶颈,提出适用于NVIDIA Jetson DLA的五步部署方法,实现近零开销的多模型分层分类,可推广至各类检测分类边缘流水线。

中文摘要 AI 辅助

边缘部署的视觉系统应用于目标识别、监控、自动驾驶和无人机领域,需要分层推理流水线,其中检测模型识别感兴趣的对象,下游分类器提供细粒度属性分析。在GPU上运行所有模型会产生串行瓶颈,随着流水线阶段增加,这会限制实时吞吐量。现代边缘片上系统(SoC)将GPU与专用神经加速器(NPU、DLA)配对,支持并发执行,但由于严格的算子约束、量化不兼容性以及无文档的端到端流水线,在这些加速器上部署自定义模型仍不现实。我们以NVIDIA Jetson DLA核心为代表性平台,提出了一种用于分类骨干网络的五步零GPU fallback DLA INT8部署方法,包括架构适配、手动动态范围变通方案以恢复TensorRT的隐式量化(从隐式量化的75%准确率恢复至94.0%),用于显式量化前的快速流水线验证、量化感知训练、用于DLA编译的ONNX图手术,以及并发GPU检测/DLA分类推理流水线。我们记录了九个工程约束及其根本原因分析和可推广的解决方案。在Jetson Orin NX上,与GPU对象检测器并行在DLA上运行的双头人物属性分类器的验证表明,流水线开销接近零(1080p下,检测器单独运行的帧率为13.3,本方法为12.5),且双DLA扩展无额外成本。该方法与骨干网络无关,可推广到任何检测-分类边缘流水线。

英文摘要

Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest and downstream classifiers provide fine-grained attribute analysis. Running all models on the GPU creates a serial bottleneck that limits real-time throughput as pipeline stages grow. Modern edge SoCs pair GPUs with dedicated neural accelerators (NPUs, DLAs) capable of concurrent execution, yet deploying custom models on these accelerators remains impractical due to strict operator constraints, quantization incompatibilities, and an undocumented end-to-end pipeline. We target NVIDIA Jetson DLA cores as the representative platform. We present a five-step methodology for zero GPU fallback DLA INT8 deployment of classification backbones, comprising architecture adaptation, manual dynamic range workaround to rescue TensorRT's implicit quantization (recovering 94.0% accuracy from implicit quantization's 75%) for rapid pipeline validation before explicit quantization, quantization-aware training, ONNX graph surgery for DLA compilation, and a concurrent GPU-detection/DLA-classification inference pipeline. We document nine engineering constraints with root-cause analysis and generalizable solutions. Validation on a dual-head person attribute classifier running on DLA alongside a GPU object detector on a Jetson Orin NX demonstrates near-zero pipeline overhead (12.5 vs. 13.3~FPS detector-only at 1080p), with dual-DLA scaling at no additional cost. The methodology is backbone-agnostic and generalizes to any detection-classification edge pipeline.

发表机构

  • Newspace Research and Technologies(纽斯佩斯研究与技术公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑