arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.22714cs.CVcs.AI

用于嵌入式汽车系统的优化RetinaNet架构的实时语义分割

Real-Time Semantic Segmentation with Optimized RetinaNet Architectures for Embedded Automotive Systems

Sai Sidharth D

首次发表
浏览论文内容

中文总结 AI 辅助

针对嵌入式汽车系统对实时感知的需求,提出Opt-RetinaSeg架构,通过优化主干、FPN及引入分割头解决不平衡问题,经三阶段优化,在相关数据集和硬件上实现高速、小模型且高精度的实时语义分割。

中文摘要 AI 辅助

实时感知是高级驾驶辅助系统(ADAS)和自动驾驶车辆的基本要求,但嵌入式汽车平台对计算、内存和功耗有严格限制。本文提出一种源自RetinaNet检测框架的优化语义分割架构,适用于密集像素级预测,并针对资源受限的嵌入式硬件进行定制。所提出的Opt-RetinaSeg架构用混合轻量级特征提取器取代标准ResNet-50主干,重组特征金字塔网络(FPN)以减少冗余多尺度计算,并引入由焦点损失启发的类平衡引导的紧凑分割头来解决道路场景中常见的严重前景-背景不平衡问题。我们进一步应用由结构化通道剪枝、训练后INT8量化和来自高容量教师网络的知识蒸馏组成的三阶段优化管道。在Cityscapes和BDD100K数据集上进行评估,并部署在NVIDIA Jetson Xavier NX和高通QCS610汽车SoC上,所提出的模型在70.4 FPS时实现了73.9%的平均交并比(mIoU),相对于ResNet-50基线,推理速度提高了7.4倍,模型大小减少了4倍,精度下降不到3%。这些结果表明,经过系统优化的源自RetinaNet的架构是嵌入式汽车感知管道中实时语义分割的可行候选方案。

英文摘要

Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded automotive platforms impose severe constraints on compute, memory, and power. This paper presents an optimized semantic segmentation architecture derived from the RetinaNet detection framework, adapted for dense pixel-wise prediction and tailored for deployment on resource-constrained embedded hardware. The proposed architecture, termed Opt-RetinaSeg, replaces the standard ResNet-50 backbone with a hybrid lightweight feature extractor, restructures the Feature Pyramid Network (FPN) to reduce redundant multi-scale computation, and introduces a compact segmentation head guided by focal-loss-inspired class balancing to address the severe foreground-background imbalance common in road scenes. We further apply a three-stage optimization pipeline consisting of structured channel pruning, post-training INT8 quantization, and knowledge distillation from a high-capacity teacher network. Evaluated on the Cityscapes and BDD100K datasets and deployed on an NVIDIA Jetson Xavier NX and a Qualcomm QCS610 automotive SoC, the proposed model achieves 73.9% mIoU at 70.4 FPS, representing a 7.4x inference speedup and a 4x reduction in model size relative to the ResNet-50 baseline, with less than 3% accuracy degradation. These results indicate that RetinaNet-derived architectures, when systematically optimized, are viable candidates for real-time semantic segmentation in embedded automotive perception pipelines

↑