Integrating large language models (LLMs) into automatic speech recognition (ASR) has become a mainstream paradigm in recent years. Although existing LLM-based ASR models demonstrate impressive performance on public benchmarks, their training remains predominantly data-driven, leaving key practical challenges insufficiently addressed -- particularly limited downward scalability in resource-constrained deployments and hallucinations under acoustically challenging conditions. To address these issues, we present NIM4-ASR, a production-oriented LLM-based ASR framework optimized for both efficiency and robustness. Grounded in a principled delineation of functional roles between the encoder and the LLM, we redesign the multi-stage training paradigm to align each module with its intended capability boundary. Specifically, we reformulate the pre-training architecture and objective to mitigate the modality gap and improve parameter efficiency; introduce an iterative asynchronous SFT stage to preserve acoustic fidelity and constrain representation drift; and design an ASR-specialized reinforcement learning stage to further enhance recognition quality and robustness. We additionally incorporate a suite of production-oriented optimizations, including robustness under noisy and silent conditions, real-time streaming inference, and hotword customization via retrieval-augmented generation (RAG). Experiments show that NIM4-ASR achieves state-of-the-art performance on multiple public benchmarks with merely 2.3B parameters, while substantially outperforming larger-scale competitors on internal benchmarks -- particularly in entity-intensive real-world scenarios. NIM4-ASR further supports million-scale hotword customization via RAG with sub-millisecond retrieval latency, enabling efficient adaptation to emerging entities and personalized user requirements.
"Parallel Training Considered Harmful?": Comparing series-parallel and parallel feedforward network training
并行训练是否有害?:比较系列-并行与并行前馈网络训练
Antônio H. Ribeiro, Luis A. Aguirre
机构
*
Department of Electronic Engineering at Universidade Federal de Minas Gerais (UFMG) - Av. Ant\ o nio Carlos 6627, 31270-901, Belo Horizonte, MG, Brazil
Neural network models for dynamic systems can be trained either in parallel or in series-parallel configurations. Influenced by early arguments, several papers justify the choice of series-parallel rather than parallel configuration claiming it has a lower computational cost, better stability properties during training and provides more accurate results. Other published results, on the other hand, defend parallel training as being more robust and capable of yielding more accu- rate long-term predictions. The main contribution of this paper is to present a study comparing both methods under the same unified framework. We focus on three aspects: i) robustness of the estimation in the presence of noise; ii) computational cost; and, iii) convergence. A unifying mathematical framework and simulation studies show situations where each training method provides better validation results, being parallel training better in what is believed to be more realistic scenarios. An example using measured data seems to reinforce such claim. We also show, with a novel complexity analysis and numerical examples, that both methods have similar computational cost, being series series-parallel training, however, more amenable to parallelization. Some informal discussion about stability and convergence properties is presented and explored in the examples.
SparseStreet: Sparse Gaussian Splatting for Real-Time Street Scene Simulation
SparseStreet: 用于实时街景模拟的稀疏高斯泼溅
Qingpo Wuwu, Xiaobao Wei, Peng Chen, Nan Huang, Zhongyu Zhao, Hao Wang, Ming Lu, Ningning Ma, Shanghang Zhang
机构
*
Peking University(北京大学)
;
Chinese Academy of Sciences(中国科学院)
;
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
;
Autonomous Driving Development, NIO(蔚来自动驾驶开发)
While 3D Gaussian Splatting has shown promising results in street scene reconstruction, existing methods require massive numbers of Gaussian primitives to capture fine details, leading to prohibitive storage costs and slow rendering speeds. We observe that dynamic objects (e.g., vehicles and pedestrians) demand high-fidelity representations to maintain temporal consistency, while static background regions often contain substantial redundancy. Motivated by this, we propose SparseStreet, a general compression framework specifically designed for street scenes. First, we introduce a node-based learnable pruning strategy that systematically removes low-contributing Gaussian primitives while preserving visually critical regions. Second, after the scene representation stabilizes, we apply background compression, further reducing redundancy in static regions. Our method effectively preserves the geometry and appearance of dynamic objects while significantly reducing the total number of Gaussian primitives. Extensive experiments on the Waymo and nuScenes demonstrate that SparseStreet achieves up to 80% compression ratio with minimal quality degradation, enabling resource-efficient, high-fidelity dynamic scene reconstruction. Project website: https://sparsestreet.github.io/.
机构
*
College of Electronic and Information Engineering, Tongji University(同济大学电子与信息学院)
;
Department of Control Science and Engineering, Harbin Institute of Technology(控制科学与工程系,哈尔滨工业大学)
;
Department of Vehicle Control System and Software Development, NIO(车辆控制系统与软件开发部,蔚来汽车)
;
School of Computer Science and Technology, Tongji University(计算机科学与技术学院,同济大学)
;
Key Laboratory of Embedded System and Service Computing (Ministry of Education), Tongji University(嵌入式系统与服务计算重点实验室(教育部),同济大学)
Reliable driving scene parsing is a fundamental capability for autonomous vehicles operating in open and dynamic driving environments. However, adapting perception models to new deployment domains remains challenging because pixel-level annotations are expensive to obtain, while source-domain data are often inaccessible due to privacy, security, or ownership constraints. Existing source-free unsupervised domain adaptation methods typically rely on a single pre-trained source model, which makes the adapted perception system vulnerable to source-specific biases and limits its robustness under diverse road layouts, illumination conditions, weather patterns, and traffic conditions. This article presents an unsupervised collaborative domain adaptation (UCDA) framework for driving scene parsing in a source-free setting, which transfers complementary knowledge from multiple pre-trained source models to a unified target model without accessing any original source samples. To compare predictions from independently trained models, UCDA constructs a class-level prototype memory bank and estimates cross-model prediction reliability through prototype similarity, reducing the effect of inconsistent confidence scales across source models. Based on the resulting complementary supervision, UCDA adopts a two-stage transfer strategy: multiple source models are first refined on unlabeled target-domain driving data through collaborative optimization with positive and negative consistency constraints, and their validated expertise is then distilled into a single deployable target model. Comprehensive evaluations on public driving-scene datasets and real-world data collected from an autonomous vehicle platform demonstrate that UCDA effectively consolidates complementary multi-source knowledge, improving target-domain scene parsing reliability and generalization across diverse driving environments.
An Approach for Thyroid Nodule Analysis Using Thermographic Images
使用热成像图像进行甲状腺结节分析的方法
J. R. González, É. O. Rodrigues, C. P. Damião, C. A. P. Fontes, A. C. Silva, A. C. Paiva, H. Li, C. Du, A. Conci
机构
*
Computer Science Department, Universidade Federal Fluminense(联邦弗里蒙特大学计算机科学系)
;
Radiology Department, Hospital Universitário Antônio Pedro (HUAP)(安东尼奥佩德罗大学医院放射科)
;
Applied Computation Group NCA-UFMA, Universidade Federal do Maranhão(马兰舍大学应用计算组NCA-UFMA)
据预测,到2030年,甲状腺癌将成为女性中第二常见的癌症类型,男性中第三常见。一般来说,早期检测癌症可提高个体生存机会。热成像是一种诊断工具,越来越多地用于检测癌症和异常,包括甲状腺异常。已有多种方法被提出用于分割和检测热成像图中的热区域,从而检测这些图像中存在的可疑组织。众所周知,医学诊断会产生大量信息。因此,医生必须在短时间内全面分析和评估这些信息,这在大多数情况下是不可行的。在这项工作中,我们对热成像进行了全面综述,重点关注甲状腺分析。我们提出了图像采集协议和甲状腺图像的自主配准方法。我们还对图像数据进行了分析,包括特征提取、图像处理以及一种可能的健康或非健康患者分类方法。总之,这项工作提出了在我们大学医院检测肿瘤的试点项目,这是支持我们内分泌科预防性医疗行动的一部分。经过一些未来调整后,该项目将提交给弗鲁米嫩塞联邦大学安东尼奥·佩德罗大学医院(HUAP-UFF)的伦理与研究委员会以及巴西卫生部伦理委员会审批,项目名称为:评估热成像在HUAP-UFF患者甲状腺结节诊断辅助中的重要性(葡萄牙语:Avaliação da importância da termografia no auxílio à investigação diagnóstica de nódulos tireoidianos em pacientes acompanhados no HUAP-UFF)。
英文摘要
Thyroid cancer is said to be the second most common type of cancer in female individuals and the third in males by 2030, according to projections. In general, detecting cancer in its early stages improves the chance of survival of the individual. Thermography is a diagnostic tool that has been increasingly used to detect cancer and abnormalities, including that of thyroid. Various methods to segment and detect hot regions in thermograms and, consequently, to detect suspicious tissues present in these images have been proposed. It is well known that medical diagnosis yields a great deal of information. Thus, physicians have to comprehensively analyse and evaluate this information in a short period of time, which is infeasible in most cases. In this work, we perform a general review of thermography , focusing on the thyroid analysis. We propose protocols for image acquisiton and an autonomous registration for thyroid images. We also perform analyses of the image data, which include feature extraction, image processing, and a possible approach for classification of healthy or unhealthy patients. In summary, this work presents a pilot project for detection of tumors in our university hospital, which is part of an effort to support preventive medical actions in our endocrinology department. Under some future adjustments, this project will be submitted for approval by the ethics and research committee of Hospital Universitário Antonio Pedro at Universidade Federal Fluminense (HUAP-UFF) and to the Brazilian Ministry of Health Ethical committee under the name: Evaluation of the importance of thermography to aid diagnosis of thyroid nodules of patients in HUAP-UFF (in Portuguese: Avaliação da importância da termografia no auxílio à investigação diagnóstica de nódulos tireoidianos em pacientes acompanhados no HUAP-UFF).
Two-Stream Interactive Joint Learning of Scene Parsing and Geometric Vision Tasks
双流交互式场景解析与几何视觉任务联合学习
Guanfeng Tang, Hongbo Zhao, Ziwei Long, Jiayao Li, Bohong Xiao, Wei Ye, Hanli Wang, Rui Fan
机构
*
College of Electronic and Information Engineering, Tongji University(同济大学电子信息学院)
;
Shanghai Research Institute for Intelligent Autonomous Systems, Tongji University(同济大学智能自主系统研究所)
;
College of Computer Science and Technology, Tongji University(同济大学计算机科学与技术学院)
;
Department of Vehicle Control System and Software Development, NIO(蔚来汽车车辆控制系统与软件开发部门)
;
Department of Automotive Engineering, Jilin University(吉林大学汽车工程学院)
;
Key Laboratory of Embedded System and Service Computing (Ministry of Education), Tongji University(同济大学嵌入式系统与服务计算重点实验室)
;
College of Electronic and Information Engineering, Shanghai Institute of Intelligent Science and Technology(上海智能科学与技术研究院电子信息学院)
;
Shanghai Key Laboratory of Intelligent Autonomous Systems(上海智能自主系统重点实验室)
Inspired by the human visual system, which operates on two parallel yet interactive streams for contextual and spatial understanding, this article presents Two Interactive Streams (TwInS), a novel bio-inspired joint learning framework capable of simultaneously performing scene parsing and geometric vision tasks. TwInS adopts a unified, general-purpose architecture in which multi-level contextual features from the scene parsing stream are infused into the geometric vision stream to guide its iterative refinement. In the reverse direction, decoded geometric features are projected into the contextual feature space for selective heterogeneous feature fusion via a novel cross-task adapter, which leverages rich cross-view geometric cues to enhance scene parsing. To eliminate the dependence on costly human-annotated correspondence ground truth, TwInS is further equipped with a tailored semi-supervised training strategy, which unleashes the potential of large-scale multi-view data and enables continuous self-evolution without requiring ground-truth correspondences. Extensive experiments conducted on three public datasets validate the effectiveness of TwInS's core components and demonstrate its superior performance over existing state-of-the-art approaches. The source code will be made publicly available upon publication.
Localization is a critical technology in autonomous driving, encompassing both topological localization, which identifies the most similar map keyframe to the current observation, and metric localization, which provides precise spatial coordinates. Conventional methods typically address these tasks independently, rely on single-camera setups, and often require additional 3D semantic or pose priors, while lacking mechanisms to quantify the confidence of localization results, making them less feasible for real industrial applications. In this paper, we propose VVLoc, a unified pipeline that employs a single neural network to concurrently achieve topological and metric vehicle localization using multi-camera system. VVLoc first evaluates the geo-proximity between visual observations, then estimates their relative metric poses using a matching strategy, while also providing a confidence measure. Additionally, the training process for VVLoc is highly efficient, requiring only pairs of visual data and corresponding ground-truth poses, eliminating the need for complex supplementary data. We evaluate VVLoc not only on the publicly available datasets, but also on a more challenging self-collected dataset, demonstrating its ability to deliver state-of-the-art localization accuracy across a wide range of localization tasks.
RAVE: End-to-end Hierarchical Visual Localization with Rasterized and Vectorized HD map
Jinyu Miao, Tuopu Wen, Kun Jiang, Kangan Qian, Zheng Fu, Yunlong Wang, Zhihuang Zhang, Mengmeng Yang, Jin Huang, Zhihua Zhong, Diange Yang
机构
*
School of Vehicle and Mobility, and State Key Laboratory of Intelligent Green Vehicle and Mobility, Tsinghua University(车辆与移动系统学院,智能绿色车辆与移动系统国家重点实验室,清华大学)
;
NIO Inc.(蔚来汽车公司)
;
Qcraft Inc.(Qcraft公司)
;
Chinese Academy of Engineering(中国工程院)
Comments16 pages, 10 figures, 6 tables
详情
英文摘要
Accurate localization serves as an important component in autonomous driving systems. Traditional rule-based localization involves many standalone modules, which is theoretically fragile and requires costly hyperparameter tuning, therefore sacrificing the accuracy and generalization. In this paper, we propose an end-to-end visual localization system, RAVE, in which the surrounding images are associated with the HD map data to estimate pose. To ensure high-quality observations for localization, a low-rank flow-based prior fusion module (FLORA) is developed to incorporate misaligned map prior into the perceived BEV features. Pursuing a balance among efficiency, interpretability, and accuracy, a hierarchical localization module is proposed, which efficiently estimates poses through a decoupled BEV neural matching-based pose solver (DEMA) using rasterized HD map, and then refines the estimation through a Transformer-based pose regressor (POET) using vectorized HD map. The experimental results demonstrate that our method can perform robust and accurate localization under varying environmental conditions while running efficiently.