arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于早期检测和跟踪隐藏动态对象的预测性音频表示

Predictive audio representations for early detection and tracking of hidden dynamic objects

Katerina Vinciguerra, Moritz Brandes, Danilo Hollosi, Letizia Marchegiani

arXiv 2609.13595首次发表:更新:

发表机构

University of Parma; Fraunhofer Institute for Digital Media Technology IDMT(帕尔马大学; 弗劳恩霍夫数字媒体技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一个基于JEPA自监督预训练和LSTM多任务微调的音频系统,同时估计隐藏车辆的数量、类型和到达方向,在非视距多智能体场景中优于现有方法。

AI 中文摘要

预测潜在危险是安全的核心。预测其他交通参与者的存在是危险预测的核心。被遮挡的交通参与者对检测系统构成挑战,因为它们可能过晚才变得可见,留给自动驾驶车辆太少的时间来以稳健且安全的方式识别、规划并采取相应行动。先前的研究证明,听觉感知具有全向性且不受视野限制,能够为早期发现不同的道路使用者提供基本线索,即使这些使用者被其他车辆或基础设施隐藏。然而,这些贡献仅处理只有一辆车存在的场景,并且它们要么识别车辆类型,要么估计其到达方向。在本工作中,我们向前推进,提出了一个多任务系统,该系统同时估计存在的车辆数量、其类型及其到达方向。我们的方法论贡献是一个两阶段流水线:一个受联合嵌入预测架构(JEPA)启发的自监督预训练阶段,直接应用于多通道原始波形,随后是使用双向长短期记忆网络(LSTM)和三个分类头的监督多任务微调。预训练阶段在没有标签的情况下训练编码器,促使其从过去的上下文预测未来音频片段的潜在表示。为了训练和测试我们的框架,由于没有公开可用的合适数据集,我们收集了一个专门的数据集,涵盖多个交通参与者同时运行的非视距(NLOS)场景。实验评估表明,我们的方法优于现有技术水平;同时证明我们的设计选择使模型能够学习稳健的表示,这些表示可以迁移到未见过的驾驶场景,保持合理且稳定的性能。

英文摘要

Predicting potential dangers is core to safety. Forecasting the presence of other traffic agents is core to danger prediction. Occluded traffic agents challenge detection systems as they might become visible too late, leaving the autonomous vehicle too little time to identify, plan and act accordingly in a robust and safe way. Previous works proved that auditory perception, being omnidirectional and not constrained by a field-of-view, provides fundamental cues for early spotting of different road users, even when hidden by other vehicles or infrastructures. Yet, those contributions deal with scenarios with only one vehicle present, and they either identify the type of vehicle or estimate its direction of arrival. In this work, we move forward, and propose a multi-task system which, simultaneously, estimates the number of vehicles present, their type, and their direction of arrival. Our methodological contribution is a two-stage pipeline: a self-supervised pre-training stage inspired by the Joint- Embedding Predictive Architecture (JEPA) applied directly to multichannel raw waveforms, followed by supervised multi-task fine-tuning with a bidirectional LSTM and three classification heads. The pre-training stage trains the encoder without labels, pushing it to predict the latent representation of a future audio segment from its past context. To train and test our framework, since no suitable dataset was publicly available, we collected an ad-hoc one covering Non-Line-Of-Sight scenarios with multiple traffic agents simultaneously operating. Experimental evaluation shows that our method outperforms the state of the art; it also proves that our design choice allows the model to learn robust representations, which can be transferred to an unseen driving scenario, maintaining reasonable and stable performance.

Comments8 pages, 2 tables, 6 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑